Back to skill

Security audit

YouTube Transcript Generator

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly matches its YouTube transcript purpose, but its bundled script has a real argument-handling flaw that can run unintended local Python code if invoked with a crafted option.

Review this skill before installing or running it. It appears intended to generate YouTube transcripts, but the bundled script should be patched to validate the timestamps option and pass dynamic values to Python as arguments rather than interpolating them into source code. Avoid letting untrusted text control the script arguments, and be aware it writes transcript files locally.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/get_transcript.sh:81
Finding

Arbitrary Python Code Execution Through Unsafe Timestamps Argument Interpolation

Content
View full analysis

Vulnerability Details

File Location: scripts/get_transcript.sh, lines 81-86
Vulnerability Type: Python code injection
Risk Level: High

Vulnerable Code

bash
# Clean VTT/SRT to plain text
python3 -c "
import re, sys

timestamps_mode = '$TIMESTAMPS' == 'timestamps'

with open('$SUBTITLE_FILE', 'r', encoding='utf-8') as f:

Technical Analysis

The fourth command-line argument is assigned to TIMESTAMPS and then interpolated directly into source code passed to python3 -c. Although the shell variable appears between single quotation marks in the generated Python statement, those quotation marks are part of the Python source rather than a shell-level security boundary.

An attacker can include a single quote and Python statement delimiters in the fourth argument. This closes the intended Python string literal and inserts arbitrary Python statements. The remainder of the original statement can then be neutralized with a Python comment.

For example, a fourth argument shaped like the following would cause an observable file write:

text
'; __import__("pathlib").Path("/tmp/transcript-injection-poc").write_text("executed"); #

This produces Python source equivalent to:

python
timestamps_mode = ''; __import__("pathlib").Path("/tmp/transcript-injection-poc").write_text("executed"); #' == 'timestamps'

The same primitive can invoke os.system, subprocess, or native Python APIs to execute commands and access files. The vulnerability does not require shell metacharacters to survive shell evaluation because the malicious value is introduced through normal argument expansion into the Python program.

The subtitle path is also embedded into Python source using the same unsafe pattern. Although the reviewed script normally derives that path from a private temporary directory and a fixed yt-dlp output template, it should still be passed as data rather than interpolated into sourc ...[truncated 1906 chars]

Remediation
View remediation

Remediation Suggestions

Never construct executable Python source by interpolating command-line arguments. Pass all dynamic values through sys.argv or environment variables and validate options before invoking Python.

A safer implementation is:

bash
case "$TIMESTAMPS" in
  ""|timestamps)
    ;;
  *)
    echo "ERROR: Fourth argument must be 'timestamps' or empty." >&2
    exit 2
    ;;
esac

python3 - "$TIMESTAMPS" "$SUBTITLE_FILE" > "$OUTPUT" <<'PY'
import re
import sys

timestamps_mode = sys.argv[1] == "timestamps"
subtitle_file = sys.argv[2]

with open(subtitle_file, "r", encoding="utf-8") as f:
    content = f.read()

# Continue transcript processing here.
PY

This hardening separates code from data, prevents quotation characters in arguments from changing Python syntax, and safely handles unusual subtitle paths. Additional input validation should restrict the timestamps mode to the documented values and reject unexpected extra options.

Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (5)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
85% confidence
Finding

The skill invokes a local shell script and its documented behavior includes writing a transcript file to disk, but the manifest declares no tool scope or permissions. This creates a transparency and containment problem: an agent or user may invoke the skill without realizing it needs file access and local execution capability, increasing the chance of unintended file reads/writes or broader execution than expected.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The description uses broad trigger phrases like 'transcribe,' 'get transcript,' and 'extract text,' which can overlap with many ordinary user requests. Over-broad activation can cause the wrong skill to run unexpectedly, leading to unintended network access, local file creation, or execution of the bundled script in contexts where the user did not specifically request this tool.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The instructions state that the script tries English manual subtitles first and then English auto-generated subtitles before trying other available languages. This imposes a language preference in the skill behavior without stating user choice or a justified locale-specific constraint.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill states in the body that it saves output to transcript_VIDEO_ID.txt, but the description does not warn users up front that local files will be created. This can surprise users, cause unintentional persistence of potentially sensitive transcript content, and complicate cleanup or privacy expectations on shared systems.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The script redirects generated transcript content into "$OUTPUT", which may point to an existing file, but there is no confirmation prompt or explicit warning about overwriting local data in the script comments or user-facing output. Because this operation modifies the filesystem, a brief disclosure or overwrite safeguard would improve user awareness.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.