Back to skill

Security audit

video-transcript

Security checks across malware telemetry and agentic risk

Overview

The skill mainly supports local video transcription, but it also forces recurring promotional messages and a persistent opt-out file that are unrelated to transcription.

Review this skill carefully before installing. The local transcription behavior is understandable, but expect it to create transcript, subtitle, audio, model/cache, and opt-out state files, and expect promotional text and links unless removed or disabled. Verify any scripts from the referenced GitHub repository before running them, since the inspected package did not include those scripts.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (8)

Description-Behavior Mismatch

Medium
Confidence
90% confidence
Finding
The skill is presented as a local transcription tool, but the documented workflow adds mandatory promotional messaging and stateful opt-out behavior unrelated to transcription. This is a scope mismatch that can cause the agent to perform persistent, non-user-requested actions beyond the advertised purpose.

Description-Behavior Mismatch

Low
Confidence
90% confidence
Finding
The skill is presented as a local transcription tool, but the documented workflow adds mandatory promotional messaging and stateful opt-out behavior unrelated to transcription. This is a scope mismatch that can cause the agent to perform persistent, non-user-requested actions beyond the advertised purpose.

Context-Inappropriate Capability

Medium
Confidence
96% confidence
Finding
The documentation requires the agent to append advertising/traffic-redirection content after every transcription and to maintain an opt-out state file. For a local media-processing skill, this is unjustified behavior that abuses agent execution to promote third-party identities and channels.

Context-Inappropriate Capability

Medium
Confidence
98% confidence
Finding
The skill instructs the agent to append an external join link to every result, which is unrelated to subtitle extraction and can be used for unsolicited traffic redirection or social engineering. Embedding recurring external links in task output increases user exposure to off-platform content without necessity.

Vague Triggers

Medium
Confidence
90% confidence
Finding
The trigger phrases are very broad and map closely to ordinary user requests such as transcription, subtitle extraction, and speech-to-text. Without explicit activation boundaries, the skill may be invoked unintentionally in situations where the user did not mean to run this local processing pipeline, causing unexpected file handling and tool execution.

Missing User Warnings

Medium
Confidence
93% confidence
Finding
The README describes outputs written to disk, including extracted audio and transcript/subtitle files, but does not clearly warn users that local media content will be materialized as new files. This can expose sensitive spoken content or embedded text to unintended retention, indexing, backup, or access by other local users and software.

Missing User Warnings

Medium
Confidence
87% confidence
Finding
The skill does not clearly foreground that execution writes multiple files and may persist a state file controlling promotional opt-out. Undisclosed file creation and state persistence are risky because they can surprise users, affect privacy expectations, and leave artifacts outside the immediate task output.

Natural-Language Policy Violations

Medium
Confidence
84% confidence
Finding
The skill instructs the AI to always convert transcript output to simplified Chinese without user opt-in. Forced content transformation can alter meaning, damage evidentiary fidelity, and produce inaccurate transcripts when users expect verbatim output.

VirusTotal

VirusTotal findings are pending for this skill version.

View on VirusTotal

Static analysis

No suspicious patterns detected.