Back to skill

Security audit

youtube-apify-transcript

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly does what it says, but a URL-handling flaw can let its cache read or write JSON files outside the intended cache folder.

Review before installing. Use this only with trusted YouTube URLs and the documented default Apify actor, keep the Apify token out of repositories and shared shell history, and prefer a temporary environment variable or secret manager. The cache path validation should be fixed before using this skill in workflows that process untrusted URLs or run with access to important JSON files.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/fetch_transcript.py:61
Finding

Path Traversal Through Insufficient YouTube Video ID Validation

Content
View full analysis

Vulnerability Details

File Location: scripts/fetch_transcript.py:61-64, scripts/fetch_transcript.py:97-119
Vulnerability Type: Path traversal and arbitrary JSON file access
Risk Level: High

Vulnerable Code

python
# Handle youtu.be short links
if "youtu.be" in url:
    path = urlparse(url).path
    return path.lstrip("/").split("?")[0]

The resulting value is subsequently used directly as part of a filesystem path:

python
def get_cache_path(video_id):
    """Get the cache file path for a video ID."""
    return CACHE_DIR / f"{video_id}.json"


def load_from_cache(video_id):
    """Load transcript from cache. Returns None if not cached."""
    cache_path = get_cache_path(video_id)
    if cache_path.exists():
        try:
            with open(cache_path, "r", encoding="utf-8") as f:
                return json.load(f)
        except (json.JSONDecodeError, IOError):
            return None
    return None


def write_cache_data(video_id, cache_data):
    """Write cache data to disk."""
    ensure_cache_dir()
    cache_path = get_cache_path(video_id)
    with open(cache_path, "w", encoding="utf-8") as f:
        json.dump(cache_data, f, ensure_ascii=False, indent=2)

Technical Analysis

The short-URL branch checks whether the untrusted input contains the substring youtu.be, rather than verifying that the parsed hostname is exactly an approved YouTube hostname. It then returns the entire URL path without enforcing the expected 11-character video ID format.

Consequently, path components such as ../ can become part of video_id. The cache path is formed using:

python
CACHE_DIR / f"{video_id}.json"

pathlib does not automatically confine this path to CACHE_DIR. Traversal components can therefore resolve to a JSON file outside the intended cache directory.

The vulnerable value reaches both read and write operations. An external J ...[truncated 2017 chars]

Remediation
View remediation

Remediation Suggestions

  1. Parse the URL before making any domain decision and compare the normalized hostname against an explicit allowlist:

    python
    ALLOWED_HOSTS = {
        "youtu.be",
        "youtube.com",
        "www.youtube.com",
        "m.youtube.com",
    }
    
  2. Extract only the expected path segment from short URLs.

  3. Validate every extracted ID, regardless of URL format, against the canonical format:

    python
    VIDEO_ID_RE = re.compile(r"^[A-Za-z0-9_-]{11}$")
    
    def validate_video_id(value):
        return value if VIDEO_ID_RE.fullmatch(value) else None
    
  4. Add defense-in-depth confinement before all cache operations:

    python
    def get_cache_path(video_id):
        if not VIDEO_ID_RE.fullmatch(video_id):
            raise ValueError("Invalid YouTube video ID")
    
        cache_root = CACHE_DIR.resolve()
        cache_path = (cache_root / f"{video_id}.json").resolve()
    
        if cache_path.parent != cache_root:
            raise ValueError("Cache path escapes cache directory")
    
        return cache_path
    
  5. Use restrictive cache permissions where supported, such as a user-only cache directory and files created with mode 0600.

  6. Add tests covering traversal strings, misleading hostnames, encoded separators, extra path segments, empty IDs, and IDs longer or shorter than 11 characters.

T08 · Insecure Dependencies

Note
Location
README.md:37
Finding

Unpinned Third-Party Dependency Installation

Content
View full analysis

Vulnerability Details

File Location: README.md:37-40; also documented in SKILL.md:19, SKILL.md:93-96, SKILL.md:221, and scripts/fetch_transcript.py:29-32
Vulnerability Type: Unpinned runtime dependency and non-reproducible installation
Risk Level: Low

Vulnerable Code

bash
# 1. Set your API token
export APIFY_API_TOKEN="apify_api_YOUR_TOKEN"

# 2. Install the Python dependency (OpenClaw has no pip installer kind)
python3 -m pip install requests

The script repeats the unconstrained installation instruction:

python
try:
    import requests
except ImportError:
    print("Error: 'requests' library not installed.", file=sys.stderr)
    print("Install with: pip install requests", file=sys.stderr)
    sys.exit(1)

Technical Analysis

The project instructs users to install requests without a pinned version, lock file, or package hash. Installation therefore resolves whichever release is current at installation time.

No typographical error, suspicious package index, or known malicious package is present in the audited files. The package name is the legitimate requests package. The risk arises from non-reproducible dependency resolution: a future compromised, incompatible, or unexpectedly changed upstream release would be selected automatically.

Because Python package installation and package imports may execute package-controlled code, compromise of the selected distribution could expose the same privileges, environment variables, and files available to the user running the Skill.

Attack Path

  1. A user follows the documented python3 -m pip install requests command.
  2. Pip resolves the newest package version available from its configured package index.
  3. If that release or the configured index is compromised, malicious installation or runtime code is placed in the Python environment.
  4. The Skill imports requests while APIFY_API_TOKEN is present in ...[truncated 936 chars]
Remediation
View remediation

Remediation Suggestions

  1. Pin an audited requests release using an exact version constraint.

  2. Prefer a committed requirements or lock file rather than installation instructions that resolve dependencies dynamically.

  3. Record package hashes and install with hash verification, for example:

    text
    requests==AUDITED_VERSION --hash=sha256:EXPECTED_HASH
    
    bash
    python3 -m pip install --require-hashes -r requirements.txt
    
  4. Include and hash-pin applicable transitive dependencies so the complete environment is reproducible.

  5. Install dependencies in a dedicated virtual environment rather than a global or privileged Python environment.

  6. Avoid running pip or the Skill as root.

  7. Keep the pinned dependency set under periodic vulnerability review and update it through an explicit tested release process.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Rogue AgentSelf-Modification, Session Persistence
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (8)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 98)May include surrounding context.

Python dependency (once)

python3 -m pip install requests

Or use .env file (never commit this!)

echo 'APIFY_API_TOKEN=apify_api_YOUR_TOKEN_HERE' >> .env

text

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 99)May include surrounding context.

Python dependency (once)

python3 -m pip install requests

Or use .env file (never commit this!)

echo 'APIFY_API_TOKEN=apify_api_YOUR_TOKEN_HERE' >> .env

text

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill documents capabilities to read environment variables, write cache/output files, and make network requests, but it does not declare any explicit tool scope or permission boundaries. That increases the chance an agent framework will grant broader-than-necessary access, making misuse of the Apify token or local file writes harder to constrain and audit.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 87)May include surrounding context.

md
## Setup

1. Create a free Apify account: https://apify.com/
2. Get your API token: https://console.apify.com/account/integrations
3. Set the environment variable and install `requests` yourself. OpenClaw has no `kind: pip` installer (allowed kinds: brew, node, go, uv, download), so this is a manual step.

Session Persistence

Medium
Category
Rogue Agent
Confidence
90% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 92)May include surrounding context.

  1. Set the environment variable and install requests yourself. OpenClaw has no kind: pip installer (allowed kinds: brew, node, go, uv, download), so this is a manual step.
bash
# Add to ~/.bashrc or ~/.zshrc
export APIFY_API_TOKEN="apify_api_YOUR_TOKEN_HERE"

# Python dependency (once)

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/fetch_transcript.py (reported line 37)May include surrounding context.

python
DEFAULT_APIFY_ACTOR_ID = "topaz_sharingan~Youtube-Transcript-Scraper-1"
APIFY_API_BASE = "https://api.apify.com/v2"
CACHE_DIR = Path(os.environ.get("YT_TRANSCRIPT_CACHE_DIR", os.path.join(os.path.dirname(os.path.dirname(os.path.abspath(__file__))), ".cache")))

ACTOR_ALIASES = {

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/fetch_transcript.py (reported line 259)May include surrounding context.

python
try:
        # Start the run
        response = requests.post(
            run_url,
            headers=headers,
            params=params,

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The setup instructions tell users to place a live API token in shell startup files or a .env file, but they do not clearly warn about exposure through shell history, shared home directories, backups, logs, or accidental repository inclusion. While common documentation practice, this is still a security weakness because it normalizes long-lived secret storage without enough operational guidance.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/fetch_transcript.py:53

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
SKILL.md:88