Back to skill

Security audit

Book Video Generator En

Security checks for vulnerabilities and agentic risk

Overview

The skill appears to make book-summary videos as advertised, but it automatically installs unpinned Python packages and sends user content to external AI/TTS services, so it needs review before installation.

Install only after reviewing the provider choices. Prefer a dedicated virtual environment, preinstall pinned dependencies yourself, and avoid letting the scripts auto-install packages during a normal run. Do not submit confidential book data, unpublished manuscripts, private branding, or sensitive prompts unless the selected search, image, and TTS providers are approved for that data. Keep SD_WEBUI_URL pointed at a trusted local service if using the local backend.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
scripts/generate_audio.py:218
Finding

Automatic Installation of Unpinned Third-Party Dependencies

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
Findings (48)

Tainted flow: 'req' from os.environ.get (line 389, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/generate_image.py (reported line 164)May include surrounding context.

python
req = urllib.request.Request(url, data=data, headers=headers, method="POST")
    try:
        with urllib.request.urlopen(req, timeout=300) as resp:
            result = json.loads(resp.read())
    except urllib.error.HTTPError as e:
        err_body = e.read().decode()[:500]

Tainted flow: 'req' from os.environ.get (line 389, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/generate_image.py (reported line 293)May include surrounding context.

python
req = urllib.request.Request(url, data=data, headers=headers, method="POST")
    try:
        with urllib.request.urlopen(req, timeout=300) as resp:
            result = json.loads(resp.read())
    except urllib.error.HTTPError as e:
        err_body = e.read().decode()[:500]

Tainted flow: 'req' from os.environ.get (line 389, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/generate_image.py (reported line 327)May include surrounding context.

python
}).encode()

    req = urllib.request.Request(url, data=data, headers=headers, method="POST")
    with urllib.request.urlopen(req) as resp:
        result = json.loads(resp.read())

    image_data = base64.b64decode(result["data"][0]["b64_json"])

Tainted flow: 'req' from os.environ.get (line 389, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/generate_image.py (reported line 362)May include surrounding context.

python
}).encode()

    req = urllib.request.Request(url, data=data, headers=headers, method="POST")
    with urllib.request.urlopen(req) as resp:
        result = json.loads(resp.read())

    image_data = base64.b64decode(result["data"][0]["b64_json"])

Tainted flow: 'req' from os.environ.get (line 389, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

The local Stable Diffusion endpoint is built from the SD_WEBUI_URL environment variable without validation, so a caller can redirect requests to an arbitrary host rather than a trusted localhost service. In agent environments, that can cause prompts and generated content requests to be sent to attacker-controlled infrastructure under the guise of a local backend.

Content

Scanner excerpt · scripts/generate_image.py (reported line 390)May include surrounding context.

python
}).encode()

    req = urllib.request.Request(url, data=data, headers={"Content-Type": "application/json"}, method="POST")
    with urllib.request.urlopen(req) as resp:
        result = json.loads(resp.read())

    image_data = base64.b64decode(result["images"][0])

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · .gitignore (reported line 48)May include surrounding context.

text
*.skill

# Environment variables
.env
.env.local

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · .gitignore (reported line 49)May include surrounding context.

text
# Environment variables
.env
.env.local

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 87)May include surrounding context.

md
| Engine | Credential | Timestamps | Notes |
|--------|-----------|-----------|-------|
| Volcano Engine TTS (default) | `VOLC_TTS_API_KEY` | Estimated from audio duration | Doubao Speech 2.0, best English naturalness, commercial use, get API Key from [Volcano console](https://console.volcengine.com/speech/new) |
| edge-tts (fallback) | none | Native WordBoundary | Microsoft free TTS, works out of the box, auto-fallback when no credential |

Set the Volcano credential:

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 377)May include surrounding context.

md
The original Coze workflow file is at `references/workflow-original.yaml`, a full chain of 30+ nodes:

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 403)May include surrounding context.

md
The original Coze workflow file is at `references/workflow-original.yaml`, a full chain of 30+ nodes:

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Skill reads from agent configuration directories (.claude/, .codex/, .gemini/). These directories may contain API keys, personal settings, and other credentials that the skill has no legitimate need to access.

Content

Scanner excerpt · SKILL.md (reported line 358)May include surrounding context.

bash
# 1. 开启 Skills 功能(config.toml)
cat >> ~/.codex/config.toml << 'EOF'
[features]
skills = true
EOF

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Skill reads from agent configuration directories (.claude/, .codex/, .gemini/). These directories may contain API keys, personal settings, and other credentials that the skill has no legitimate need to access.

Content

Scanner excerpt · references/CROSS_PLATFORM.md (reported line 81)May include surrounding context.

bash
# 1. 开启 Skills 功能(config.toml)
cat >> ~/.codex/config.toml << 'EOF'
[features]
skills = true
EOF

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
85% confidence
Finding

Skill reads from agent configuration directories (.claude/, .codex/, .gemini/). These directories may contain API keys, personal settings, and other credentials that the skill has no legitimate need to access.

Content

Scanner excerpt · references/CROSS_PLATFORM.md (reported line 81)May include surrounding context.

bash
# 1. 开启 Skills 功能(config.toml)
cat >> ~/.codex/config.toml << 'EOF'
[features]
skills = true
EOF

External Script Fetching

High
Category
Supply Chain
Confidence
97% confidence
Finding

The guide suggests piping content fetched from an external website directly into Python via curl ... | python3 -c "...", which normalizes an unsafe pattern of executing data derived from the network. Even though the placeholder ... implies incompleteness, this example encourages command constructions where untrusted remote content can influence code execution, creating a clear path to arbitrary code execution if copied or adapted unsafely.

Content

Scanner excerpt · references/CROSS_PLATFORM.md (reported line 99)May include surrounding context.

方案 A — Shell 命令搜索(免安装):

bash
curl -s "https://www.google.com/search?q=book+title+author+summary" | python3 -c "..."

方案 B — 安装搜索 MCP 插件。

os.system() or os exec-family call

High
Category
Dangerous Code Execution
Confidence
95% confidence
Finding

The script executes a shell command to install a missing dependency at runtime using os.system. Even though sys.executable is usually controlled by the local environment, invoking the shell for package installation creates an unnecessary command-execution path, can run unreviewed code from PyPI, and may behave dangerously in privileged or automated environments.

Content

Scanner excerpt · scripts/generate_audio.py (reported line 221)May include surrounding context.

python
import edge_tts
    except ImportError:
        print("Installing edge-tts...")
        os.system(f"{sys.executable} -m pip install edge-tts -q")
        import edge_tts

    os.makedirs(os.path.dirname(output_path) or ".", exist_ok=True)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill is presented as an 'English Version' and states that it produces the entire output 'entirely in English.' This is a natural-language locale constraint without any indication that users can choose another language or explicitly opt in, which matches the policy concern for forced language selection.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The markdown explicitly states that all script, narration, on-screen text, and prompts are fully in English. This is a natural-language locale policy constraint presented as mandatory behavior, and the file does not offer the user a language choice or document a justified region-specific requirement.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The workflow directs the agent to use third-party web search, image generation, and TTS providers, and elsewhere references API-based services and credentials, but it does not give a clear up-front privacy warning that user-supplied book titles, prompts, captions, and possibly derived content may be transmitted externally. This creates a real data-handling and consent risk, especially in enterprise or regulated environments where users may assume processing is local.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
85% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · SKILL.md (reported line 360)May include surrounding context.

md
# 1. Enable Skills (if not yet)
echo '[features]\nskills = true' >> ~/.codex/config.toml

# 2. Copy skill dir
cp -r ~/.workbuddy/skills/book-video-generator-en ~/.codex/skills/

# 3. Restart Codex CLI

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file repeatedly specifies that the skill is the English edition and, in this section, explicitly constrains TTS, subtitles, and fonts to English-only rendering. Because no user choice or opt-in is provided for language/locale behavior, this is a natural-language locale policy concern under the stated rule.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The file states that all output is in English, and later prompt fields also require English output for generated content. This imposes a language policy without offering the user a language choice or documenting a justified locale constraint.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The prompt explicitly instructs the model to use tools to search for real information about the book, but gives no user-facing notice that external/network retrieval may occur or what data will be sent. In an agent skill, this can expose user-supplied inputs such as book titles, authors, or surrounding context to third-party services without informed consent, creating privacy and compliance risk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The top-level description states the workflow generates a Chinese-language video format, and multiple downstream prompts are written to require Chinese output. This imposes a specific language/locale behavior on all users without offering a choice or documenting a justified regional constraint.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The workflow sends generated and user-derived content, including text, audio, and images, to multiple external plugins and hosted asset URLs without any privacy or network-use disclosure. In this context, book inputs may seem low sensitivity, but account identifiers, generated media, and potentially user-provided names or branding still leave the local workflow boundary and may be retained or processed by third parties.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The workflow creates and saves an external video draft through plugins, changing external project state without any visible confirmation, warning, or consent gate. In an agent setting, silent state-changing actions can surprise users, create unauthorized artifacts under a linked account, and make abuse or accidental misuse harder to detect.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/generate_image.py:144