Back to skill

Security audit

Baoyu YouTube Transcript

Security checks for vulnerabilities and agentic risk

Overview

This skill mostly does what it says, but it can read browser session cookies and dynamically acquire external tools without clear user approval boundaries.

Review this skill before installing if you do not want agents to dynamically fetch runtimes or fallback tools. Do not set YOUTUBE_TRANSCRIPT_COOKIES_FROM_BROWSER unless you intentionally want yt-dlp to access your browser profile cookies for YouTube, because those cookies can represent an authenticated session. Expect transcript text, metadata, descriptions, and thumbnails to be saved locally in the configured output directory.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
Findings (16)

YARA rule 'info_stealer': Information stealer patterns (credential harvesting, browser data theft) [malware]

High
Category
YARA Match
Confidence
78% confidence
Finding

This match is not evidence of malware or an information stealer by itself, but it does identify a genuinely sensitive behavior: accessing browser cookies to authenticate fallback downloads. In the context of an agent skill, encouraging cookie extraction without strong warnings or approval controls can expose session credentials and make the capability significantly more dangerous than ordinary transcript retrieval.

Content

Scanner excerpt · SKILL.md (reported line 82)May include surrounding context.

md
d transcripts | |
| `--exclude-manually-created` | Skip manually created transcripts | |
| `--refresh` | Force re-fetch, ignore cached data | |
| `-o, --output <path>` | Save to specific file path | auto-generated |
| `--output-dir <dir>` | Base output directory | `youtube-transcript` |

## Optional Environment Variables

| Variable | Description |
|----------|-------------|
| `YOUTUBE_TRANSCRIPT_COOKIES_FROM_BROWSER` | Passed to `yt-dlp --cookies-from-browser` during fallback, e.g. `chrome`, `safari`, `firefox`, or `chrome:Profile 1` |

## Input Formats

Accepts any of these as video input:
- Full URL: `https://www.youtube.com/watch?v=dQw4w9WgXcQ`
- Short URL: `https://youtu.be/dQw4w9WgXcQ`
- Embed URL: `https://www.youtube.com/embed/dQw4w9WgXcQ`
- Shorts URL: `https://www.youtube.com/shorts/dQw4w9WgXcQ`
- Video ID: `dQw4w9WgXcQ`

## Output Formats

| Format | Extension | Description |
|--------|-----------|-------------|
| `text` | `.md` | Markdown with frontmatter (incl. `descriptio

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill explicitly uses network access to YouTube and environment variables, but it declares no corresponding tool scope or permission boundaries. That creates an authorization gap where the agent may invoke broader capabilities than users or platform policy expect, reducing transparency and weakening containment.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill caches transcript text, metadata, descriptions, and images to disk by default, but the description does not clearly warn about this persistent local storage. Users may unknowingly retain potentially sensitive or copyrighted content, increasing privacy, compliance, and data handling risk.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
92% confidence
Finding

Using npx -y bun without a pinned version allows execution of whatever package version is current at runtime, which is a classic supply-chain risk. If the upstream package is compromised or changed incompatibly, the skill may execute unreviewed code on the host.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The browser-cookie fallback can access authenticated browser session data via --cookies-from-browser, but the documentation does not clearly warn about the privacy and account-security implications. Even if intended for legitimate access, this materially increases the risk of mishandling sensitive session tokens or broadening access to protected content.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

Telling the agent to make yt-dlp available on its own introduces autonomous tool acquisition and possible package installation outside the declared purpose. That behavior increases supply-chain and execution risk because the agent may fetch and run external software without a clear approval boundary.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill goes beyond transcript retrieval by instructing the agent to spawn a separate AI sub-agent and process local files with additional prompts. This expands the trust boundary and creates extra opportunities for prompt injection, unintended data exposure, or uncontrolled secondary actions unrelated to the core file-download purpose.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The code hard-codes English locale settings in request parameters and headers, including hl=en and Accept-Language: en-US,en;q=0.9. This appears to force a specific language/locale rather than offering user choice or documenting a justified locale restriction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The InnerTube request context sets hl to en and gl to US for all requests. This enforces a specific language and region behavior without any visible opt-in or configurable alternative.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This skill invokes external programs (yt-dlp, uvx, python3 -m yt_dlp) to satisfy requests, which expands the trust boundary from simple HTTP fetching into local code execution. That is risky because the behavior now depends on whatever executable is present in the environment, its version, and its plugins/components, increasing attack surface beyond the stated transcript/thumbnail purpose.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

The fallback command uvx --from yt-dlp yt-dlp pulls and executes an unpinned package at runtime, creating a supply-chain execution path outside the skill's direct codebase. If the upstream package, dependency resolution, or transport path is compromised, this skill could execute attacker-controlled code when attempting transcript retrieval.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Reading YOUTUBE_TRANSCRIPT_COOKIES_FROM_BROWSER enables the skill to direct yt-dlp to access browser cookies, which is a credential-adjacent capability not obviously required for ordinary public transcript retrieval. In a shared agent environment, this can turn a simple content-fetching tool into one that accesses authenticated session material and potentially private or age-restricted content.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

downloadCoverImage writes fetched bytes directly to the provided outputPath without any validation or confinement. If an attacker can influence that path through higher layers, this permits arbitrary file overwrite within the agent's filesystem permissions, which can corrupt files or place attacker-chosen content in sensitive locations.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The workflow says to overwrite the .md file with processed transcript content, which is a destructive file-write operation. Although this may be part of the feature flow, the markdown does not clearly warn that the original raw transcript file will be replaced unless backed up elsewhere.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The file explicitly requires using the same language as the transcription for the title and table of contents. This imposes a language/locale behavior without indicating that the user can choose a different output language, which matches the policy's language-choice concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

Transcript retrieval uses a fixed Accept-Language header of en-US,en;q=0.9. This is another hard-coded locale choice that may override user expectations or organizational language policy.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dangerous_exec

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
scripts/youtube.ts:293