Back to skill

Security audit

Video Insight

Security checks for vulnerabilities and agentic risk

Overview

This video transcript skill is mostly purpose-aligned, but it automatically uses browser cookies and has unsafe network/credential handling that users should review before installing.

Review before installing. Avoid using this skill on private or account-restricted videos unless the browser-cookie fallback is removed or made explicitly opt-in. Do not run --summarize with sensitive transcripts unless the LLM endpoint is trusted and uses safe transport, and avoid setting LLM_API_URL alongside OPENCLAW_GATEWAY_TOKEN. Prefer an isolated virtual environment and pinned dependencies before setup.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (5)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/utils.py:49
Finding

Unvalidated Bilibili URLs Enable Arbitrary Network Requests and Automatic Browser-Cookie Use

Content
View full analysis
str: """Detect video platform from URL.""" if "bilibili.com" in url or "b23.tv" in url: return "bilibili" elif "youtube.com" in url or "youtu.be" in url: return "youtube" return "unknown" ``` ```python if "b23.tv" in url: try: import requests r = requests.head(url, allow_redirects=True, timeout=10) url = r.url except Exception: pass ``` ```python base_cmd = [ "yt-dlp", "--no-check-certificates", "--retries", retries, "--fragment-retries", frag_retries, "-f", "bestvideo[height<=720]+bestaudio/best[height<=720]", "--merge-output-format", "mp4", "-o", video_path, url, ] try: subprocess.run(base_cmd, check=True, timeout=yt_dlp_timeout, capture_output=True) except subprocess.CalledProcessError: progress(" ⚠️ Retrying with browser cookies...") cookie_cmd = [ "yt-dlp", "--cookies-from-browser", "chrome", "--no-check-certificates", "--retries", retries, "--fragment-retries", frag_retries, "-f", "bestvideo[height<=720]+bestaudio/best[height<=720]", "--merge-output-format", "mp4", "-o", video_path, url, ] subprocess.run(cookie_cmd, check=True, timeout=yt_dlp_timeout, capture_output=True) ``` ### Technical Analysis Platform detection is based on substring matching against the entire URL rather than parsing and validating the hostname. An attacker can therefore supply a URL whose query, path, user-information component, or attacker-controlled hostname merely contains `bilibili.com` or `b23.tv`. The URL is then passed unchanged to `yt-dlp`. When `b ...[truncated 2217 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/llm.py:110
Finding

OpenClaw Gateway Token Can Be Sent to an Attacker-Controlled LLM Endpoint

Content
View full analysis
` and the full summarization prompt to that endpoint. 7. The attacker ...[truncated 537 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/llm.py:55
Finding

Custom LLM Configuration Permits Cleartext Transmission of API Keys and Full Transcripts

Content
View full analysis
Optional[str]: """Call an OpenAI-compatible chat API.""" try: import requests headers = {"Content-Type": "application/json"} if api_key: headers["Authorization"] = f"Bearer {api_key}" response = requests.post( api_url, headers=headers, json={ "model": model, "messages": [{"role": "user", "content": prompt}], "max_tokens": 4000, }, timeout=timeout, ) if response.status_code == 200: return response.json()["choices"][0]["message"]["content"] ``` ```python env_url = os.environ.get("LLM_API_URL") env_key = os.environ.get("LLM_API_KEY") env_model = os.environ.get("LLM_MODEL", "gpt-4o-mini") if env_url and env_key: progress(f" 🔑 LLM: {env_url}") result = _call_llm(env_url, env_key, env_model, prompt) if result: return result ``` ### Technical Analysis `LLM_API_URL` is accepted without validating its scheme or destination. If it uses `http://`, the bearer API key and complete prompt are transmitted without transport encryption. The prompt contains the full transcript without truncation, together with video metadata. The code also relies on the default redirect behavior of `requests`. If an endpoint redirects the request, the final destination is not explicitly checked against an approved host policy. Although `requests` may strip authorization across some cross-origin redirects, the implementation should not rely on implicit library behavior to protect sensitive credentials and content. ### Attack Path 1. A user or wrapper c ...[truncated 797 chars]
Remediation
View remediation

T01 · Skill Instruction Hijacking

Warning
Location
scripts/llm.py:17
Finding

Untrusted Video Transcripts Can Inject Instructions into LLM Summaries

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
setup.sh:126
Finding

Setup Installs Unpinned Dependencies and Can Modify System Python

Content
View full analysis
/dev/null || "$PIP" install "$@" else "$PIP" install --quiet --break-system-packages "$@" 2>/dev/null || \ "$PIP" install --quiet --user "$@" 2>/dev/null || \ "$PIP" install "$@" fi } # ── Core dependencies ── log "Installing core dependencies..." pip_install yt-dlp youtube-transcript-api innertube requests ``` ```bash pip_install faster-whisper ``` ### Technical Analysis The setup script installs packages using mutable package names without exact versions or cryptographic hashes. Each setup run can therefore retrieve different package versions than those reviewed during the audit. If a future release or transitive dependency is compromised, malicious installation or runtime code can execute with the privileges of the user running setup. If virtual-environment creation fails, the script falls back to `--break-system-packages`, user installation, or ordinary pip installation. This can modify the system or user Python environment and affect unrelated applications, exceeding the isolation expected from a Skill installer. No evidence was found that the named dependencies are intentionally malicious. The issue is the absence of reproducible dependency controls and the unsafe system-environment fallback. ### Attack Path 1. A user follows `SKILL.md` and runs `bash setup.sh`. 2. The script requests the latest versions satisfying each unpinned package name. 3. A compromised future package release or transitive dependency is downloaded. 4. Package installation or imported runtime code executes under the user’s privileges. 5. If virtual-environment creation failed, installation may alter the broader syst ...[truncated 568 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Output HandlingUnvalidated Output Injection, Cross-Context Output, Unbounded Output
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
Findings (33)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill description emphasizes transcript extraction, but the documented behavior includes downloading broader video content and extracting keyframes, which materially expands data collection and processing beyond what a user may expect. This mismatch is security-relevant because agents and users rely on metadata to make trust and consent decisions; understated behavior can lead to unintended network usage, storage of media-derived artifacts, and privacy exposure.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The skill description emphasizes transcript extraction, but the documented behavior includes downloading broader video content and extracting keyframes, which materially expands data collection and processing beyond what a user may expect. This mismatch is security-relevant because agents and users rely on metadata to make trust and consent decisions; understated behavior can lead to unintended network usage, storage of media-derived artifacts, and privacy exposure.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
93% confidence
Finding

Using '--no-check-certificates' disables TLS certificate validation for yt-dlp downloads, making the connection vulnerable to man-in-the-middle interference. In a tool that retrieves remote media for later processing, this can allow tampered content or metadata to be injected and undermines transport security.

Content

Scanner excerpt · scripts/bilibili.py (reported line 139)May include surrounding context.

python
progress("  📥 Downloading video...")
        base_cmd = [
            "yt-dlp",
            "--no-check-certificates",
            "--retries", retries,
            "--fragment-retries", frag_retries,
            "-f", "bestvideo[height<=720]+bestaudio/best[height<=720]",

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

On download failure, the code silently retries using '--cookies-from-browser chrome', which accesses authenticated browser cookie material without explicit user consent at the point of use. In an agent skill context, this is especially dangerous because it expands from transcript extraction into harvesting local browser session data to access protected content, creating privacy, account, and policy risks.

Content

No source excerpt is available for this finding.

YARA rule 'info_stealer': Information stealer patterns (credential harvesting, browser data theft) [malware]

High
Category
YARA Match
Confidence
95% confidence
Finding

The match is justified because the code reads cookies directly from Chrome via yt-dlp. While this does not look like classic malware exfiltration, it is still a real credential-access pattern: the skill reaches into a browser profile to obtain session tokens, which is inappropriate and high risk for the stated transcript-extraction purpose.

Content

Scanner excerpt · scripts/bilibili.py (reported line 153)May include surrounding context.

python
720]",
            "--merge-output-format", "mp4",
            "-o", video_path,
            url,
        ]
        try:
            subprocess.run(base_cmd, check=True, timeout=yt_dlp_timeout, capture_output=True)
        except subprocess.CalledProcessError:
            progress("  ⚠️  Retrying with browser cookies...")
            cookie_cmd = [
                "yt-dlp",
                "--cookies-from-browser", "chrome",
                "--no-check-certificates",
                "--retries", retries,
                "--fragment-retries", frag_retries,
                "-f", "bestvideo[height<=720]+bestaudio/best[height<=720]",
                "--merge-output-format", "mp4",
                "-o", video_path,
                url,
            ]
            subprocess.run(cookie_cmd, check=True, timeout=yt_dlp_timeout, capture_output=True)

        # Step 2: Extract audio
        progress("  🎵 Extracting audio...")
        subprocess.run(
            ["ffmpeg", "-y", "-i", video_

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
94% confidence
Finding

This retry path combines disabled certificate validation with browser-cookie-based authenticated requests, which is a particularly risky combination. An attacker positioned on the network could interfere with an authenticated media retrieval flow, and the code lowers transport protections exactly when it is using sensitive browser-derived session state.

Content

Scanner excerpt · scripts/bilibili.py (reported line 154)May include surrounding context.

python
cookie_cmd = [
                "yt-dlp",
                "--cookies-from-browser", "chrome",
                "--no-check-certificates",
                "--retries", retries,
                "--fragment-retries", frag_retries,
                "-f", "bestvideo[height<=720]+bestaudio/best[height<=720]",

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · scripts/llm.py (reported line 53)May include surrounding context.

python
prompt_file = skill_root / "config" / "prompt.md"
    if prompt_file.exists():
        try:
            return prompt_file.read_text(encoding="utf-8")
        except Exception:
            pass
    return DEFAULT_PROMPT

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill advertises capabilities that imply network, shell, filesystem, and environment access, but it does not declare any explicit tool scope or permission boundaries. In an agent environment, this weakens policy enforcement and increases the chance the skill is invoked with broader privileges than intended, especially because setup and runtime commands interact with external services and local disk.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The trigger list contains broad phrases like general video summary/transcript requests, increasing the chance the skill is auto-selected for loosely related prompts. Over-broad triggering can cause unintended network access, transcript retrieval, local caching, or keyframe extraction when the user did not explicitly ask to invoke this tool.

Content

No source excerpt is available for this finding.

Unbounded Output

Medium
Category
Output Handling
Confidence
88% confidence
Finding

Returning full transcripts without truncation creates unbounded output that can exhaust agent context windows, memory, disk, logs, or downstream processing limits. In multi-step agent systems, excessively large outputs can also amplify prompt-injection exposure embedded in transcripts and increase operational cost or denial-of-service risk.

Content

Scanner excerpt · SKILL.md (reported line 56)May include surrounding context.

md
"title": "Video Title",
    "channel": "Channel Name",
    "duration_seconds": 212,
    "transcript": "Full transcript text without truncation...",
    "transcript_with_timestamps": "[0.0-3.2] First segment\n[3.2-6.5] Second...",
    "frames": [{"file": "/tmp/.../frame_001.jpg", "time_sec": 30}],
    "cached": false

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill permanently caches transcripts and metadata on disk, but this retention is not clearly surfaced in the user-facing description. Transcripts may contain sensitive spoken content, and persistent local storage increases the risk of later unauthorized access, data leakage between users, or retention beyond user expectations.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest describes a transcript-first skill with optional AI summarization, while this module docstring explicitly states "download + whisper transcription + optional keyframes" and the implementation downloads the video and can extract JPEG frames. Keyframe/image extraction is a broader media-processing capability not reflected in the stated skill description.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/bilibili.py (reported line 23)May include surrounding context.

python
"""Get Bilibili video title via yt-dlp --print title."""
    retries = int(get_setting("yt_dlp_retries", 3, settings=SETTINGS))
    try:
        result = subprocess.run(
            ["yt-dlp", "--print", "title", "--retries", str(retries), "--no-warnings", url],
            capture_output=True, text=True, timeout=30,
        )

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/bilibili.py (reported line 38)May include surrounding context.

python
"""Get Bilibili video duration via yt-dlp."""
    retries = int(get_setting("yt_dlp_retries", 3, settings=SETTINGS))
    try:
        result = subprocess.run(
            ["yt-dlp", "--print", "duration", "--retries", str(retries), "--no-warnings", url],
            capture_output=True, text=True, timeout=30,
        )

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/bilibili.py (reported line 53)May include surrounding context.

python
"""Get Bilibili uploader name via yt-dlp."""
    retries = int(get_setting("yt_dlp_retries", 3, settings=SETTINGS))
    try:
        result = subprocess.run(
            ["yt-dlp", "--print", "uploader", "--retries", str(retries), "--no-warnings", url],
            capture_output=True, text=True, timeout=30,
        )

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/bilibili.py (reported line 148)May include surrounding context.

python
url,
        ]
        try:
            subprocess.run(base_cmd, check=True, timeout=yt_dlp_timeout, capture_output=True)
        except subprocess.CalledProcessError:
            progress("  ⚠️  Retrying with browser cookies...")
            cookie_cmd = [

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The retry path introduces credentialed browser state into the workflow without any user-facing warning or consent prompt. Hidden use of browser cookies materially changes the trust boundary and can expose session data or cause the tool to access private/account-scoped content unexpectedly.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/bilibili.py (reported line 162)May include surrounding context.

python
"-o", video_path,
                url,
            ]
            subprocess.run(cookie_cmd, check=True, timeout=yt_dlp_timeout, capture_output=True)

        # Step 2: Extract audio
        progress("  🎵 Extracting audio...")

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/bilibili.py (reported line 166)May include surrounding context.

python
# Step 2: Extract audio
        progress("  🎵 Extracting audio...")
        subprocess.run(
            ["ffmpeg", "-y", "-i", video_path, "-vn", "-acodec", "libmp3lame", "-q:a", "2", audio_path],
            check=True, timeout=ffmpeg_timeout, capture_output=True,
        )

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/bilibili.py (reported line 214)May include surrounding context.

python
if extract_frames:
            progress(f"  🖼️  Extracting keyframes (every {frame_interval}s)...")
            os.makedirs(frames_dir, exist_ok=True)
            subprocess.run(
                [
                    "ffmpeg", "-y", "-i", video_path if os.path.exists(video_path) else audio_path,
                    "-vf", f"fps=1/{frame_interval}", "-q:v", "2",

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · scripts/cli.py (reported line 45)May include surrounding context.

python
parser.add_argument("--daily", action="store_true", help="Daily batch mode (with --config)")

    # Output control
    parser.add_argument("--output", "-o", help="Write JSON to file instead of stdout")
    parser.add_argument("--quiet", "-q", action="store_true", help="Suppress stderr progress")
    parser.add_argument("--pretty", action="store_true", default=True, help="Pretty-print JSON (default)")
    parser.add_argument("--compact", action="store_true", help="Compact JSON output")

External Transmission

Medium
Category
Data Exfiltration
Confidence
89% confidence
Finding

This network call transmits prompt content, including the full transcript, to a configured API endpoint. Because the endpoint is environment-controlled and may be external, this is a genuine outbound data transfer risk in the context of potentially sensitive transcript content.

Content

Scanner excerpt · scripts/llm.py (reported line 67)May include surrounding context.

python
if api_key:
            headers["Authorization"] = f"Bearer {api_key}"

        response = requests.post(
            api_url,
            headers=headers,
            json={

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The generate_summary docstring states, 'By default the agent does its own summarization,' which contradicts this function's actual behavior. The implementation only sends the transcript to an OpenAI-compatible API endpoint or OpenClaw Gateway if environment variables are present, and otherwise returns None with 'Skipping summary.'

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The code sends the full transcript to an LLM backend, which may be a remote third-party endpoint via LLM_API_URL, without any technical guardrail or explicit disclosure at the transmission point. Video transcripts can contain sensitive, copyrighted, or user-provided content, so forwarding them externally creates a data exposure risk, even if summarization is opt-in.

Content

No source excerpt is available for this finding.

Ssd 1

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The transcript is embedded directly into the prompt as plain instruction text, so adversarial video content can inject prompt-level instructions that influence the model's summary. While this does not directly execute code here, it can cause unsafe, manipulated, or policy-bypassing outputs and is especially relevant because transcripts are untrusted external content.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.