Back to skill

Security audit

YouTube Transcript (yt-dlp captions)

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly extracts YouTube transcripts as advertised, but it ships unused third-party network code that conflicts with its privacy claims and should be reviewed before installation.

Review this skill before installing. It appears designed for YouTube transcript extraction, but the publisher should remove the dormant third-party provider functions or clearly document and gate them behind explicit opt-in. Only provide cookies when necessary, keep them in ~/.config/yt-transcript, and remember that transcript results are cached locally in SQLite.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/yt_transcript.py:435
Finding

Dormant Undisclosed Third-Party Requests and Provider-Controlled URL Fetching

Content
View full analysis
tuple[str, str, list[dict[str, Any]]]: """Last-resort fallback via a third-party transcript provider.""" api = "https://yt-to-text.com/api/v1/Subtitles" body = json.dumps({"video_id": video_id}).encode("utf-8") headers = { "User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120 Safari/537.36", "Accept": "application/json", "Content-Type": "application/json", "x-source": "tubetranscript", "x-app-version": "1.0", } req = urllib.request.Request(api, data=body, headers=headers, method="POST") try: with urllib.request.urlopen(req, timeout=30) as resp: raw = resp.read().decode("utf-8", errors="ignore") except urllib.error.HTTPError as e: msg = e.read().decode("utf-8", errors="ignore") raise TranscriptError(f"Third-party fallback (yt-to-text) HTTP {e.code}: {msg[:200]}") except urllib.error.URLError as e: raise TranscriptError(f"Third-party fallback (yt-to-text) failed: {e}") ``` ```python def _thirdparty_downsub(video_id: str) -> tuple[str, str, list[dict[str, Any]]]: """Third-party fallback using DownSub backends.""" try: from Crypto.Cipher import AES # type: ignore from Crypto.Protocol.KDF import EVP_BytesToKey # type: ignore from Crypto.Random import get_random_bytes # type: ignore except Exception as e: raise TranscriptError( "DownSub fallback requires pycryptodome (Crypto). Install: pip install pycryptodome" ) def _b64e(b: bytes) -> str: return base64.b64encode(b).decode("ascii") def _b64url(s: str) -> str: return s.replace("+ ...[truncated 5596 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (14)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared purpose is mostly aligned with the main user-facing behavior: extracting YouTube transcripts, preferring manual captions, supporting timestamps, and caching results in SQLite. However, the description is materially incomplete in a few security-relevant ways. First, the implementation is not limited to yt-dlp; it also directly fetches YouTube watch pages and calls the youtubei get_transcript endpoint. Second, it can load local Netscape-format YouTube/Google cookies and use them in requests, including constructing SAPISIDHASH authorization headers for logged-in access, which is a meaningful resource/access capability not mentioned in the declared permissions or description. Third, the code includes implemented third-party fallback functions that send video identifiers/URLs to external services (yt-to-text, DownSub, Noteey). Even though the current main path says third-party fallbacks are disabled and does not call them, the supplied code chunk still contains that undeclared capability. Because these are more than incidental implementation details and touch network destinations and authentication resources not disclosed by the description, this should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

These functions add unjustified capability to transmit user-requested video identifiers and retrieve transcript content from unrelated third-party services such as yt-to-text, DownSub, and Noteey. In a transcript skill whose stated purpose is YouTube caption extraction, this is dangerous because it silently broadens data exposure to external providers with different privacy, logging, and trust characteristics.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill advertises executable behavior involving environment variables, file access, network access, and shell execution, but does not declare an explicit tool scope such as permissions or allowed-tools. This weakens least-privilege controls and makes it harder for a host agent or reviewer to constrain what the skill may access at runtime.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The invocation description uses broad phrases like requests for transcripts, captions, subtitles, or turning a YouTube link into text, without precise trigger boundaries. Over-broad invocation criteria can cause the skill to run in contexts the user did not intend, leading to unnecessary network access, cookie use, or processing of sensitive links.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SKILL.md (reported line 85)May include surrounding context.

md
- **Never commit / publish cookies** to ClawHub.

Recommended local path (ignored by git/publish):
- `{baseDir}/cache/youtube-cookies.txt` (chmod 600)

## Notes (safety + reliability)

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/yt_transcript.py (reported line 62)May include surrounding context.

python
def _run(cmd: list[str], timeout_s: int = 60) -> str:
    try:
        p = subprocess.run(
            cmd,
            check=False,
            stdout=subprocess.PIPE,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

When a cookies file is provided, the helper builds a Cookie header from YouTube/Google cookies and attaches it to outbound HTTP requests without an in-band warning at the transmission point. That can send authenticated session material to YouTube endpoints in a way users or integrators may not fully expect, increasing privacy and account-context risk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The code hard-codes English locale preferences such as "Accept-Language: en-US,en;q=0.9", uses "hl=en" on YouTube watch URLs, defaults to English subtitle selection, and calls a fallback API with "language=en-US". This imposes a specific language/locale behavior rather than offering user choice or clearly documenting a justified regional constraint.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The file contains multiple implemented third-party transcript retrieval paths despite the skill metadata describing extraction from YouTube captions via yt-dlp. Even if currently unreachable in main flow, retaining this code creates undeclared external data-flow capability, increases attack surface, and can later be enabled accidentally or intentionally without users realizing their video identifiers may be sent to unrelated services.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
97% confidence
Finding

This code references a third-party API endpoint for fetching YouTube subtitles from Noteey, which is outside the skill's stated YouTube/yt-dlp-only scope. Such undeclared external transmission can leak video identifiers and transcript-related requests to another service with separate data retention and privacy practices.

Content

Scanner excerpt · scripts/yt_transcript.py (reported line 631)May include surrounding context.

python
"""Third-party fallback using Noteey subtitles API.

    Endpoint:
      GET https://api.noteey.com/api/v1/youtube/subtitles?url=<youtube-url>&language=<lang>

    Returns (lang, source, segments).
    """

External Transmission

Medium
Category
Data Exfiltration
Confidence
98% confidence
Finding

At this call site, the skill constructs a request to api.noteey.com containing the YouTube URL and preferred language, causing user-request data to be transmitted to a third-party service. In the context of a transcript-extraction skill advertised around yt-dlp and existing YouTube captions, that undisclosed outbound flow is a real privacy and trust concern.

Content

Scanner excerpt · scripts/yt_transcript.py (reported line 638)May include surrounding context.

python
# Noteey expects full URL
    yt_url = f"https://youtu.be/{video_id}"
    api = "https://api.noteey.com/api/v1/youtube/subtitles"
    qs = urllib.parse.urlencode({"url": yt_url, "language": "en-US"})
    raw = _http_get(api + "?" + qs, cookies=None, timeout_s=30)

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The code derives authenticated YouTube request headers, including SAPISIDHASH-based authorization, from supplied browser cookies to make account-context API requests. That expands the skill from simple public caption retrieval into authenticated session use, which can expose private account context and create privacy/security concerns if users are not clearly informed or if cookies are over-shared.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The comments say there are no third-party fallbacks in this build, but the file still contains third-party fallback implementations. This mismatch is dangerous because it can mislead reviewers and operators, hide latent capabilities during audit, and make it easier for risky external transmission paths to persist unnoticed.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The path-safety section says cookies are allowed only under ~/.config/yt-transcript/, but later recommends storing cookies at {baseDir}/cache/youtube-cookies.txt. Contradictory guidance can cause operators or calling agents to place secrets in a path the safety model claims should be disallowed, increasing the chance of accidental exposure or bypass logic errors.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.