Back to skill

Security audit

Edge TTS Voice System

Security checks for vulnerabilities and agentic risk

Overview

This voice skill has a coherent goal, but its privacy/offline claims conflict with hosted TTS behavior and its audio-file handling can expose users to code execution from crafted filenames.

Do not install this in sensitive or offline-only environments as written. Before use, require corrected privacy documentation, explicit opt-in for hosted Edge TTS, safe filename handling that avoids shell/Python source injection, pinned dependencies, and a clearer installer that avoids privileged host changes unless the user explicitly approves them.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (5)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/voice_handler.py:30
Finding

Shell and Python Code Injection Through Attacker-Controlled Audio Paths

Content
View full analysis
/dev/null" subprocess.run(cmd, shell=True, check=True) audio_file = wav_file # Transcribe with faster-whisper cmd = [ sys.executable, "-c", """ from faster_whisper import WhisperModel import sys model = WhisperModel('%s', device='cpu', compute_type='int8') segments, info = model.transcribe('%s', beam_size=5) text = ' '.join(segment.text for segment in segments) print(json.dumps({'text': text, 'language': info.language, 'probability': info.language_probability})) """ % (self.stt_model, audio_file) ] ``` ### Technical Analysis The `audio_file` value is inserted directly into a shell command enclosed only by single quotation marks. A path containing a single quote can terminate the quoted argument and introduce additional shell syntax. The resulting string is then executed with `shell=True`, giving the shell an opportunity to interpret injected metacharacters and commands. The path is subsequently interpolated into Python source passed to `python -c`. Quoting the path inside the generated Python source does not make it safe: a path containing a quote, closing parenthesis, or newline can alter the resulting Python program. The vulnerable method is exposed through `VoiceHandler.audio_to_text()` and through the module's command-line entry point. Consequently, any integration that passes externally influenced attachment paths to this method may expose the process to code execution. ### Attack Path 1. An attacker causes an audio attachment or local file to be stored under a path containing shell or Python syntax. 2. OpenClaw or another integrat ...[truncated 919 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/voice_integration.sh:38
Finding

Python Code Injection in the Shell Transcription Interface

Content
View full analysis
/dev/null ``` ### Technical Analysis The shell variable `audio_file` is expanded directly inside a single-quoted Python string in source code passed to `python -c`. Shell quoting of the overall command does not escape the value for the Python language. A filename containing a single quote and valid Python syntax can terminate the argument to `model.transcribe()`, append attacker-controlled statements, and comment out the remaining source. The preceding `-f` check only confirms that a path exists; it does not make the path safe to embed in executable source. The vulnerable function is reachable through both the `transcribe` and `process` commands. ### Attack Path 1. An attacker creates or supplies an audio file whose filename contains Python syntax. 2. A caller runs `voice_integration.sh transcribe ` or `voice_integration.sh process `. 3. The script confirms that the crafted path exists. 4. The path is expanded into the source string supplied to `python -c`. 5. Python parses the attacker-controlled portion as executable code. 6. The injected code runs with the privileges and environment of the Skill process. ### Impact Assessment An attacker can execute arbitrary Python and operating-system commands as the user running the integration script. This can expose readable OpenClaw files, cached audio, configura ...[truncated 264 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/voice_handler.py:30
Finding

Predictable Temporary Files Permit Symlink and File-Overwrite Attacks

Content
View full analysis
/dev/null" subprocess.run(cmd, shell=True, check=True) audio_file = wav_file ``` ```python if output_file is None: output_file = tempfile.mktemp(suffix=".wav", prefix="tts_") ``` ### Technical Analysis `tempfile.mktemp()` returns a candidate pathname without securely creating the file. There is a race between selecting the pathname and the later operation that creates or writes to it. A local attacker who can observe or predict the selected path may create a symbolic link or another filesystem object at that location before ffmpeg or the TTS implementation opens it. The `-y` ffmpeg option increases overwrite risk because it permits an existing destination to be replaced. The generated temporary files are also not consistently deleted, potentially retaining transcribed or generated voice data in the system temporary directory. ### Attack Path 1. The Skill selects a temporary pathname using `tempfile.mktemp()`. 2. Before ffmpeg or the TTS process opens the path, a local attacker creates a symbolic link at that pathname. 3. The link points to another file writable by the Skill's account. 4. The downstream process follows the link and writes audio output to the attacker-selected destination. 5. If the process is privileged, files available to that privileged account may be overwritten or corrupted. ### Impact Assessment The direct impact is unauthorized overwrite or corruption of files writable by the Skill account. Depending on the target file and runtime privileges, this could cause denial of service, configuration modification, or escalation through replacement ...[truncated 191 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
scripts/install.sh:71
Finding

Unpinned Network Dependencies Create a Supply-Chain Execution Risk

Content
View full analysis
Remediation
View remediation

other

Warning
Location
SKILL.md:10
Finding

False Offline and Privacy Guarantees Conceal Hosted TTS Data Processing

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (33)

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The README claims the system works 'entirely offline' while the documented stack includes Edge TTS and external model downloads, which can mislead users about privacy and network exposure. In a voice-processing skill, false offline/privacy claims are security-relevant because users may route sensitive audio through a component they believe never contacts external services.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The documentation simultaneously advertises the system as 'fully offline' and 'all processing happens locally' while identifying Edge TTS as the TTS engine, creating a material contradiction about the trust boundary. This can cause operators to deploy the skill in sensitive environments under false assumptions about confidentiality and egress.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The stated functionality appears internally inconsistent: the documentation says Edge TTS is used, but related references suggest a different or undefined TTS script and no substantiated caching or deeper OpenClaw integration. Such inconsistencies are security-relevant because they obstruct accurate review and can hide unexpected execution paths or undeclared dependencies.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
91% confidence
Finding

The stated functionality appears internally inconsistent: the documentation says Edge TTS is used, but related references suggest a different or undefined TTS script and no substantiated caching or deeper OpenClaw integration. Such inconsistencies are security-relevant because they obstruct accurate review and can hide unexpected execution paths or undeclared dependencies.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The stated functionality appears internally inconsistent: the documentation says Edge TTS is used, but related references suggest a different or undefined TTS script and no substantiated caching or deeper OpenClaw integration. Such inconsistencies are security-relevant because they obstruct accurate review and can hide unexpected execution paths or undeclared dependencies.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The file explicitly claims the system works entirely offline and that no data leaves the machine, while also stating outbound TTS uses Edge TTS, which is normally a hosted service. In a voice-processing skill, this is especially dangerous because users may send sensitive spoken content believing it remains local when it may be transmitted externally.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The documentation simultaneously makes privacy guarantees and describes a likely cloud-backed TTS component, creating a materially misleading trust signal. This is security-significant because the skill processes human speech and could expose sensitive prompts, responses, or derived personal information to third parties.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
95% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · scripts/install.sh (reported line 159)May include surrounding context.

sh
PY
    then
        log_info "✓ TTS test passed"
        rm -f "/tmp/test_install.mp3"
    else
        log_error "TTS test failed"
        return 1

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
99% confidence
Finding

This is a concrete tool-parameter abuse issue: untrusted input is interpolated into a shell command for ffmpeg. An attacker can supply a malicious filename to trigger arbitrary command execution, and in the context of an agent skill handling external files, that can lead to full compromise of the host or data exposure.

Content

Scanner excerpt · scripts/voice_handler.py (reported line 34)May include surrounding context.

python
if audio_file.endswith('.ogg'):
                wav_file = tempfile.mktemp(suffix=".wav")
                cmd = f"ffmpeg -i '{audio_file}' -ar 16000 -ac 1 '{wav_file}' -y 2>/dev/null"
                subprocess.run(cmd, shell=True, check=True)
                audio_file = wav_file
            
            # Transcribe with faster-whisper

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · README.md (reported line 65)May include surrounding context.

3. Manual Installation

bash
# Install dependencies
sudo apt-get install ffmpeg python3-pip

# Create virtual environment
python3 -m venv ~/.openclaw/tts/venv

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · README.md (reported line 164)May include surrounding context.

3. Manual Installation

bash
# Install dependencies
sudo apt-get install ffmpeg python3-pip

# Create virtual environment
python3 -m venv ~/.openclaw/tts/venv

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SKILL.md (reported line 187)May include surrounding context.

3. Manual Installation

bash
# Install dependencies
sudo apt-get install ffmpeg python3-pip

# Create virtual environment
python3 -m venv ~/.openclaw/tts/venv

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 67)May include surrounding context.

md
# Install dependencies
sudo apt-get install ffmpeg python3-pip

# Create virtual environment
python3 -m venv ~/.openclaw/tts/venv
source ~/.openclaw/tts/venv/bin/activate

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The installation instructions direct users to fetch model files from external servers without warning, despite presenting the system as offline and privacy-focused. This omission weakens informed consent and supply-chain awareness, especially for users expecting no network dependency during setup.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill documentation advertises shell-based installation and execution paths but does not declare any explicit tool scope or permissions. This weakens reviewability and containment because consumers cannot easily tell that the skill expects command execution, package installs, and system modification behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill describes automatic voice-message detection, silent transcription, AI response generation, and automatic voice replies without an explicit user-consent or notice model. In context, this increases privacy risk because users may not realize recordings are being processed and answered automatically, especially in shared or sensitive environments.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The documentation markets the voice system as local/private while also acknowledging provider-backed Edge TTS and automatic model downloads from HuggingFace. This can mislead operators into assuming stronger privacy and offline guarantees than actually exist, causing accidental transmission of data to external services or unreviewed model retrieval.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The installer performs apt-get update/install automatically, modifying the host system without confirmation or a dry-run path. In skill-install contexts this is risky because users may expect isolated setup, but instead the script changes global packages and assumes privilege escalation context, increasing supply-chain and system-integrity exposure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The installer creates and overwrites files under a default home-directory path and copies executable scripts there without prompting or backup behavior. This can unexpectedly replace an existing installation or alter runtime behavior in a path later imported by the agent, which is particularly sensitive in an agent skill ecosystem.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The installer and skill description present the system as local/private, but the install flow explicitly configures Edge TTS as a hosted service. This can mislead operators into sending text content to an external provider when they expected an offline-only workflow, creating privacy and compliance risk rather than direct code-execution risk.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/test_skill.py (reported line 81)May include surrounding context.

python
print("\nTesting ffmpeg...")
    
    try:
        result = subprocess.run(["ffmpeg", "-version"], capture_output=True, text=True)
        if result.returncode == 0:
            # Extract version
            version_line = result.stdout.split('\n')[0]

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · scripts/test_skill.py (reported line 157)May include surrounding context.

python
print("⚠ Some tests failed. See above for details.")
        print("\nCommon issues:")
        print("1. Run install.sh to install dependencies")
        print("2. Ensure ffmpeg is installed: sudo apt-get install ffmpeg")
        print("3. Check Python packages: pip install faster-whisper piper-tts soundfile")
        return 1

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

Labeling Edge TTS as 'local' without technical enforcement can mislead operators into processing sensitive voice content under false privacy assumptions. In a voice workflow, that mismatch matters because users may submit private audio and text believing no external service is involved.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
99% confidence
Finding

This call builds a shell command with an attacker-influenced file path and executes it with shell=True. A crafted audio_file containing quotes or shell metacharacters can break out of the quoted context and execute arbitrary commands under the privileges of the agent, which is especially dangerous because the skill processes external voice-message files.

Content

Scanner excerpt · scripts/voice_handler.py (reported line 34)May include surrounding context.

python
if audio_file.endswith('.ogg'):
                wav_file = tempfile.mktemp(suffix=".wav")
                cmd = f"ffmpeg -i '{audio_file}' -ar 16000 -ac 1 '{wav_file}' -y 2>/dev/null"
                subprocess.run(cmd, shell=True, check=True)
                audio_file = wav_file
            
            # Transcribe with faster-whisper

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/voice_handler.py (reported line 49)May include surrounding context.

python
""" % (self.stt_model, audio_file)
            ]
            
            result = subprocess.run(cmd, capture_output=True, text=True)
            if result.returncode == 0:
                data = json.loads(result.stdout.strip())
                return data['text'].strip()

Static analysis

No suspicious patterns detected.