Back to skill

Security audit

Video to Text (ZH)

Security checks for vulnerabilities and agentic risk

Overview

The skill’s video transcription purpose is coherent, but installation deserves review because it can automatically fetch video links, use a headless browser, and disables TLS certificate verification during downloads.

Review before installing. This skill is suitable only if you are comfortable with your agent fetching third-party video URLs, possibly opening them in headless Chromium, downloading media temporarily, and processing the audio locally. Avoid using it on sensitive or private videos unless you control the environment, and prefer a revised version that validates TLS certificates, asks before processing bare links or browser fallback, and pins dependencies.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (22)

Vague Triggers

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The trigger condition is overly broad because it mandates automatic execution whenever a user sends a video link, even without clearly requesting transcription. In context, this can cause unsolicited downloading, browser automation, and content processing on external URLs, which raises privacy, consent, and resource-consumption risks and makes prompt-trigger abuse easier.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

README 明确声明将视频链接一键转成“中文文字稿”,但没有说明这是可选项,也没有提供按原始语言或用户指定语言输出的选择。根据语言/locale 政策,强制特定语言输出而无用户选择属于自然语言策略问题。

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The README describes installation and usage but does not clearly disclose that the skill will contact third-party services, fetch remote URLs, and download media content from user-supplied links. In an agent-skill context, this can lead users to unknowingly trigger external network access, data transfer, and interaction with untrusted sites, which is especially relevant because the skill explicitly supports browser-based fallback for anti-bot sites.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill documents capabilities that imply shell, network, file read/write, and environment access, but it does not declare any explicit tool scope or permission boundaries. This increases the chance that an agent runtime grants broader access than necessary, making misuse or unexpected side effects more likely if the script or its dependencies are invoked.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill describes automatically downloading media, falling back to headless Chromium capture, and transcribing the audio, but it does not clearly disclose these data-handling behaviors to the user before execution. In this context, the absence of notice and consent is dangerous because users may not expect remote content retrieval, local media storage, or automated processing of potentially sensitive speech.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

“始终用中文回复用户”属于明确的语言/locale 强制要求,但文档没有提供用户选择语言的选项,也没有说明该限制的必要性。根据规则,未获用户选择或合理地域性说明的语言强制属于自然语言政策违规。

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
99% confidence
Finding

TLS certificate verification is explicitly disabled during short-link expansion, allowing a man-in-the-middle attacker to spoof redirects and control the resolved destination. In this skill, that is more dangerous because the resolved URL is then used to drive downstream media retrieval and browser automation against attacker-influenced targets.

Content

No source excerpt is available for this finding.

Unsafe Defaults

Medium
Category
Tool Misuse
Confidence
99% confidence
Finding

Disabling certificate validation for outbound HTTPS requests creates an unsafe default that permits interception and tampering of responses. In this workflow, an attacker could alter redirect resolution and steer the tool toward malicious or unintended media endpoints, compounding risk because the result feeds later download and parsing stages.

Content

Scanner excerpt · scripts/video_to_text.py (reported line 38)May include surrounding context.

python
"""跟随 302 展开短链接"""
    import requests
    r = requests.head(url, headers={"User-Agent": UA}, allow_redirects=True,
                      timeout=30, verify=False)
    return r.url

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/video_to_text.py (reported line 70)May include surrounding context.

python
"-o", str(audio) + ".%(ext)s", "--no-playlist",
           "--socket-timeout", "30", url]
    log("尝试 yt-dlp 下载:", url)
    r = subprocess.run(cmd, capture_output=True, text=True, timeout=600,
                       encoding="utf-8", errors="replace")
    if r.returncode != 0:
        log("yt-dlp 失败:", (r.stderr or r.stdout).strip().splitlines()[-1][:200])

File System Enumeration

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Code scans file system directories looking for sensitive files. This could be reconnaissance for credential theft.

Content

Scanner excerpt · scripts/video_to_text.py (reported line 97)May include surrounding context.

python
cands = []
    for pat in pats:
        cands += _glob.glob(pat)
    cands += _glob.glob(os.environ.get("PLAYWRIGHT_BROWSERS_PATH", "/nonexistent") + "/chromium-*/*/chrome*")
    return sorted(cands)[-1] if cands else None

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
76% confidence
Finding

This function launches headless Chromium, inspects page/network responses, and downloads media content to a local file, which affects user data and performs network activity beyond simple transcription. The script includes operational logs, but it does not clearly disclose to the user that browser automation will inspect network responses and save fetched media locally as part of fallback behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The browser context is forced to locale="zh-CN", which imposes a specific locale policy in natural-language behavior. The file does not present this as an opt-in or region-specific constraint justified to the user.

Content

No source excerpt is available for this finding.

Unsafe Defaults

Medium
Category
Tool Misuse
Confidence
99% confidence
Finding

The direct media download also disables TLS verification, allowing a network attacker to replace the downloaded audio/video with arbitrary content. Given that the file is then handed to ffmpeg and speech-processing components, this increases exposure to malicious media payloads and integrity loss of the transcription output.

Content

Scanner excerpt · scripts/video_to_text.py (reported line 171)May include surrounding context.

python
suffix = ".m4a" if audio_urls else ".mp4"
    dest = out_dir / ("source_audio" + suffix)
    log("下载媒体:", pick[:120], "...")
    r = requests.get(pick, headers=HEADERS, timeout=180, verify=False, stream=True)
    r.raise_for_status()
    with open(dest, "wb") as f:
        for chunk in r.iter_content(1 << 18):

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/video_to_text.py (reported line 183)May include surrounding context.

python
"""用 imageio-ffmpeg 的完整版 ffmpeg 统一转成 16kHz 单声道 WAV"""
    import imageio_ffmpeg
    ff = imageio_ffmpeg.get_ffmpeg_exe()
    r = subprocess.run([ff, "-y", "-i", str(src), "-vn",
                        "-acodec", "pcm_s16le", "-ar", "16000", "-ac", "1",
                        str(wav)], capture_output=True, text=True, timeout=300,
                       encoding="utf-8", errors="replace")

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The transcription function sets language="zh" and uses a Chinese initial prompt, forcing a specific language/locale in the skill's behavior. There is no user opt-in or configurable alternative, so this is a natural-language policy violation under the language-choice rule.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
96% confidence
Finding

The dependency is specified with a lower bound only, so builds may resolve to different versions over time, reducing reproducibility and making it hard to ensure a known-safe release of requests is installed. In a skill that fetches remote video content, network-facing libraries are part of the attack surface, so lack of pinning modestly increases supply-chain and patch-verification risk.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
requests>=2.31
imageio-ffmpeg>=0.5
faster-whisper>=1.0
yt-dlp>=2024.0

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
90% confidence
Finding

requests has multiple known advisories, and because the manifest does not pin a specific version, there is no way to verify whether the deployed build includes a fixed or affected release. In a skill that likely downloads or queries remote resources, uncertainty around a network library's version weakens assurance and can leave latent vulnerabilities untracked.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
93% confidence
Finding

imageio-ffmpeg is unpinned, which means the installed ffmpeg wrapper version can vary between environments and over time. Because this skill processes untrusted media from external sites, reproducible dependency selection matters for reducing exposure to unexpected regressions or vulnerable transitive behavior.

Content

Scanner excerpt · requirements.txt (reported line 2)May include surrounding context.

text
requests>=2.31
imageio-ffmpeg>=0.5
faster-whisper>=1.0
yt-dlp>=2024.0
playwright>=1.40

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
92% confidence
Finding

faster-whisper is unpinned, so deployments may silently pick different versions with different security or behavior characteristics. Since this package handles audio transcription workflows on attacker-controlled media inputs, exact version control is important for reproducibility and incident response.

Content

Scanner excerpt · requirements.txt (reported line 3)May include surrounding context.

text
requests>=2.31
imageio-ffmpeg>=0.5
faster-whisper>=1.0
yt-dlp>=2024.0
playwright>=1.40

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
97% confidence
Finding

yt-dlp is unpinned despite being a complex, high-risk component that retrieves and parses content from many third-party video platforms. In this skill's context, that makes version drift more dangerous because extractor, parser, and downloader behavior can change frequently and may expose the system to known issues or malicious upstream content-handling paths.

Content

Scanner excerpt · requirements.txt (reported line 4)May include surrounding context.

text
requests>=2.31
imageio-ffmpeg>=0.5
faster-whisper>=1.0
yt-dlp>=2024.0
playwright>=1.40

Unverifiable Dependency: yt-dlp has 16 known advisory(ies) (CVE-2023-46121 (yt-dlp Generic Extractor MITM Vulnerability via Arbitrary Proxy Injection); GHSA-3v33-3wmw-3785 (yt-dlp has dependency on potentially malicious third-party code in Douyu extract); CVE-2023-40581 ( yt-dlp on Windows vulnerable to `--exec` command injection when using `%q`) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
95% confidence
Finding

yt-dlp has a history of security advisories, and the absence of version pinning makes it impossible to determine whether the installed release is vulnerable. This is more concerning here because the skill is explicitly designed to fetch and process attacker-influenced video URLs from many public platforms, making yt-dlp a particularly exposed component.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
94% confidence
Finding

playwright is unpinned, allowing browser automation dependencies to vary across installations. Because browser automation may load untrusted pages to access video content, exact version control is important to reduce exposure to regressions and known browser-driving vulnerabilities.

Content

Scanner excerpt · requirements.txt (reported line 5)May include surrounding context.

text
imageio-ffmpeg>=0.5
faster-whisper>=1.0
yt-dlp>=2024.0
playwright>=1.40

Static analysis

Detected: suspicious.insecure_tls_verification

HTTPS certificate verification is disabled.

Warn
Code
suspicious.insecure_tls_verification
Location
scripts/video_to_text.py:38