Back to skill

Security audit

Pdf Vocab Audio

Security checks for vulnerabilities and agentic risk

Overview

The skill does what it claims, but it can send PDF-derived text to an external TTS service and writes predictable output in /tmp without clear enough privacy and filesystem safety controls.

Install only if you are comfortable with PDF-derived text being processed by edge-tts and with generated MP3s being written to /tmp. Avoid confidential PDFs unless the publisher adds an explicit privacy notice, stricter extraction, a private output directory, and pinned dependencies.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Warning
Location
pdf_vocab_audio.py:177
Finding

Predictable Shared Output Path Permits Symlink-Based File Overwrite

Content
View full analysis
Remediation
View remediation

other

Warning
Location
pdf_vocab_audio.py:68
Finding

PDF Text Is Transmitted to an External TTS Service Without Explicit Disclosure

Content
View full analysis
bool: """Generate word audio""" edge_tts_path = shutil.which("edge-tts") if not edge_tts_path: raise RuntimeError("edge-tts is not installed") cmd = [ edge_tts_path, "--voice", VOICE, f"--rate={RATE}", "--text", word, "--write-media", output_path ] result = subprocess.run( cmd, capture_output=True, text=True, timeout=30 ) ``` ```python for i, word in enumerate(words): word_file = f"{tmpdir}/word_{i}.mp3" if generate_word_audio(word, word_file): # First reading audio_segments.append(word_file) audio_segments.append(silence_file) # Second reading audio_segments.append(word_file) audio_segments.append(silence_file) ``` The extraction logic can capture the complete remainder of any line beginning with an ASCII letter: ```python match = re.match(r'^([a-zA-Z].*)$', line) if match: english = match.group(1).strip() if english and len(english) >= 2: words.append(english) ``` ### Technical Analysis The script supplies each extracted string to the `edge-tts` command-line client. Edge TTS is a network-backed synthesis mechanism, so the supplied text is transmitted to an external service to produce audio. The documentation explains that Edge TTS generates the audio, but it does not clearly state that PDF-derived content leaves the local machine, identify the external processing boundary, or request consent before transmission. The exposure is broader than the documented claim that only English words and spaces are retained. The regular expression `^([a-zA-Z].*)$` accepts any complete line beginning with an ASCII letter, ...[truncated 1545 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
SKILL.md:16
Finding

Third-Party Python Dependencies Are Installed Without Version or Integrity Pinning

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (11)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill documentation indicates capabilities to read files, write output, and invoke shell-accessible binaries (edge-tts, ffmpeg) but does not declare any explicit tool scope such as permissions or allowed-tools. This creates an authorization and review gap: an agent or platform may grant broader-than-necessary access, making misuse or accidental overreach harder to constrain, especially since the skill processes user-supplied PDFs and writes files to /tmp.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The security documentation claims all paths are temporary and non-sensitive, but the implementation reads from /root/.openclaw/media/inbound and writes to /tmp. Misleading security claims can cause operators to trust the skill with sensitive PDFs under false assumptions, increasing the chance of unintended disclosure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The script extracts text from PDFs and sends it to an external TTS program without any explicit privacy warning or consent flow. In context, PDFs may contain proprietary or personal vocabulary lists, and users may not realize content is being passed to another component that could have network behavior or logging side effects.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · pdf_vocab_audio.py (reported line 79)May include surrounding context.

python
"--text", word,
        "--write-media", output_path
    ]
    result = subprocess.run(
        cmd,
        capture_output=True,
        text=True,

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
74% confidence
Finding

Using edge-tts and ffmpeg through subprocess.run gives the skill process-execution capability. While audio generation is consistent with the overall goal, spawning external binaries is a stronger capability than the manifest describes and is not explicitly declared in scope.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · pdf_vocab_audio.py (reported line 102)May include surrounding context.

python
"-q:a", "9",
        output_path
    ]
    result = subprocess.run(cmd, capture_output=True, timeout=10)
    return result.returncode == 0

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
80% confidence
Finding

The ffmpeg concat operation consumes a list file containing file paths written without escaping or rejecting special characters such as quotes or newlines. Because output file names are derived from the PDF basename and temporary/audio paths are then written into concat_list.txt, a crafted filename could corrupt the concat manifest and potentially make ffmpeg read unintended files or fail in unsafe ways.

Content

Scanner excerpt · pdf_vocab_audio.py (reported line 131)May include surrounding context.

python
"-c", "copy",
            output_path
        ]
        result = subprocess.run(cmd, capture_output=True, timeout=60)
        return result.returncode == 0
    finally:
        if os.path.exists(list_file):

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

When no path is supplied, the script silently scans a hard-coded inbound directory and processes the newest PDF. This creates an implicit data access behavior not obvious from the skill description and could cause unintended processing of sensitive documents placed in that directory.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The markdown specifies a fixed British English male voice (en-GB-RyanNeural) and presents it as the default behavior, with no indication that users can choose another language or locale. This is a natural-language locale constraint and can violate language/locale policy when no opt-in or alternative is offered.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The docstring and all printed user-facing messages are written in Chinese, and the output filename also uses Chinese text. This imposes a specific language/locale on users without any documented choice, opt-in, or region-specific justification.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

Writing the generated MP3 to a fixed world-known location under /tmp can expose output data to other local users or processes, depending on system configuration and file permissions. It also contradicts the stated design goal of keeping paths temporary and non-sensitive.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.