Back to skill

Security audit

Truly Local Piper Multilang TTS (secure)

Security checks for vulnerabilities and agentic risk

Overview

This skill is a coherent local text-to-speech tool, with some normal setup and supply-chain risks that users should understand before installing.

Install only if you are comfortable with a one-time PyPI install into a local venv and optional HTTPS voice-model downloads. Treat 'offline' as meaning speech generation after setup; model acquisition and setup need network access. Prefer reviewing or pinning dependencies and only downloading voices you selected.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T08 · Insecure Dependencies

Warning
Location
index.js:198
Finding

Unpinned Python Dependencies Installed During Setup

Content
View full analysis
Remediation
View remediation
pathvalidate== ``` 2. Generate a locked requirements file that includes all transitive dependencies. 3. Record trusted SHA-256 hashes and install with: ```bash pip install --require-hashes -r requirements.txt ``` 4. Review and update the lock file through a controlled dependency-update process. 5. Prefer prebuilt, verified wheels and reject unexpected source builds where practical. 6. Continue requiring explicit user approval before setup, while clearly disclosing that third-party installation code runs with the Agent user's privileges. ]]>

T08 · Insecure Dependencies

Note
Location
index.js:359
Finding

Downloaded Voice Models Are Not Cryptographically Verified

Content
View full analysis
= 300 && res.statusCode < 400 && res.headers.location) { file.destroy(); try { fs.unlinkSync(dest); } catch (_) {} settled = true; downloadFile(res.headers.location, dest, redirects - 1).then(resolve, reject); return; } ``` ### Technical Analysis Voice models and their JSON metadata are downloaded over HTTPS and atomically renamed into place, but their contents are not checked against trusted hashes or digital signatures. Atomic renaming prevents partially downloaded files from being treated as complete; it does not establish artifact authenticity or integrity. In addition, redirects are accepted based only on the destination using HTTPS. There is no hostname allowlist binding the download to Hugging Face or another explicitly trusted distribution host. An attacker would need control over the upstream repository, publishing account, HTTPS redirect chain, or another trusted distribution component. This is therefore a supply-chain hardening issue rather than evidence of active malicious behavior in the skill. ### Attack P ...[truncated 1028 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
Findings (24)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · README.md (reported line 202)May include surrounding context.

Remove

bash
rm -rf ~/.openclaw/skills/local-piper-tts-multilang-secure

This removes everything: skill code, venv, and all voice models.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · SKILL.md (reported line 216)May include surrounding context.

Remove

bash
rm -rf ~/.openclaw/skills/local-piper-tts-multilang-secure

This removes everything: skill code, venv, and all voice models.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · README.md (reported line 202)May include surrounding context.

Remove

bash
rm -rf ~/.openclaw/skills/local-piper-tts-multilang-secure

This removes everything: skill code, venv, and all voice models.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · SKILL.md (reported line 216)May include surrounding context.

Remove

bash
rm -rf ~/.openclaw/skills/local-piper-tts-multilang-secure

This removes everything: skill code, venv, and all voice models.

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · SKILL.md (reported line 184)May include surrounding context.

md
4. Returns `{ removed, filesDeleted }` on success
5. If the removed voice was the user's preferred voice, ask them to pick a new one

**Never remove the last remaining voice without warning the user that TTS will stop working.**

## Changing speech speed

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · README.md (reported line 43)May include surrounding context.

bash
# Debian / Ubuntu
sudo apt install espeak-ng

# Fedora / RHEL
sudo dnf install espeak-ng

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · README.md (reported line 46)May include surrounding context.

md
sudo apt install espeak-ng

# Fedora / RHEL
sudo dnf install espeak-ng

# macOS
brew install espeak

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SKILL.md (reported line 204)May include surrounding context.

md
sudo apt install espeak-ng

# Fedora / RHEL
sudo dnf install espeak-ng

# macOS
brew install espeak

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · index.js (reported line 232)May include surrounding context.

js
sudo apt install espeak-ng

# Fedora / RHEL
sudo dnf install espeak-ng

# macOS
brew install espeak

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 195)May include surrounding context.

md
- Output filename sanitised with `path.basename()` — no directory traversal
- HTTPS-only downloads — non-HTTPS URLs and redirects are rejected
- URL path components validated against expected patterns
- Atomic downloads (write to .tmp, rename on success) — no corrupt models from interrupted downloads
- Piper installed in isolated venv — no system Python packages touched
- No credentials, no network calls during TTS (only during setup and voice downloads)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill explicitly relies on environment state (OPENCLAW_WORKSPACE, PIPER_VOICE_MODEL) and installation behavior, but it declares no permissions or allowed-tools scope. Missing capability declarations can cause the agent to invoke the skill without clear user/admin visibility into what resources it needs, increasing the chance of unintended environment access or policy bypass.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The markdown says the user may trigger this behavior with phrases like "speak faster," "too slow," or "speed it up." These are broad everyday expressions that could appear in normal conversation and the file does not provide scope limits or negative examples to clarify when the skill should act on them.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 201)May include surrounding context.

md
- otherwise: `~/.openclaw/workspace/tts/`

## Dependencies
- `python3` (3.8+) — required for `setup()` to create the venv
- `ffmpeg` — for WAV → OGG/Opus conversion
- `espeak-ng` — system library used by Piper internally; `setup()` checks for it and warns if missing.
  Install: `sudo apt install espeak-ng` (Debian/Ubuntu), `sudo dnf install espeak-ng` (Fedora),

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The manifest and metadata emphasize a local offline TTS skill, but the implementation installs packages from PyPI in setup() and later downloads voice models over HTTPS. While these actions support the TTS purpose, they materially exceed a plain 'offline' behavior claim and should be disclosed in the description as installation/provisioning requiring network access.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The setup() docstring says 'Everything stays self-contained — nothing is written outside the skill directory,' but the function invokes python3, pip, and checks for system binaries. Even if direct file writes target the skill directory, pip installation and subprocess execution are not accurately conveyed by that absolute statement, making the documentation overly contradictory to actual behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The removeVoice() function irreversibly deletes installed voice model files with fs.unlinkSync, but unlike setup() there is no confirmation prompt, warning comment, or user-facing disclosure indicating that deletion should be confirmed first. Because this operation removes user-installed assets from disk, it meets the destructive-operation criterion for missing user warnings in code files.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The script accepts a caller-supplied output path via the second argument and passes it directly to Piper's --output_file option, without restricting it to the OpenClaw workspace or validating that it stays within an approved directory. In a skill that is described as writing only into the workspace, this creates a path-traversal/arbitrary file write primitive that could overwrite files accessible to the running user, making the manifest claim materially misleading.

Content

No source excerpt is available for this finding.

Scope Creep

Low
Category
Excessive Agency
Confidence
70% confidence
Finding

Skill's behavior or capabilities extend beyond its stated purpose. Scope creep allows an agent to perform actions unrelated to its documented functionality, increasing the attack surface.

Content

Scanner excerpt · README.md (reported line 99)May include surrounding context.

await downloadVoices(['en_US-ryan-medium', 'en_US-amy-medium']);

text

Or just ask your agent: *"Download the English voices"* — it will handle everything including
playing samples so you can choose.

To see what is installed:

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

The skill markets itself as 'fully offline' and only later explains that downloadVoices() fetches models from the internet and stores them locally. That mismatch can mislead users about network behavior and persistence, weakening informed consent and increasing the chance that remote downloads occur unexpectedly in a supposedly offline workflow.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

The stated description says the skill is self-contained, offline, and extensible to any language, but does not say the code can fetch models directly from HuggingFace. That remote acquisition behavior is a meaningful operational capability beyond simple local synthesis and should be explicitly documented.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

This manifest file says the skill provides 'automatic language detection' and is 'for OpenClaw', but it does not specify what user requests or contexts should activate the skill. In a manifest, that lack of trigger specificity can lead to overly broad matching or unintended invocation because no explicit scope, examples, or exclusions are provided.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
91% confidence
Finding

Using an unpinned peer dependency of "openclaw": "*" allows installation with any available version, including future vulnerable or incompatible releases. In a skill that writes into the OpenClaw workspace, this increases supply-chain and runtime risk because the actual host API/security posture is left uncontrolled.

Content

Scanner excerpt · package.json (reported line 24)May include surrounding context.

json
},
  "dependencies": {},
  "peerDependencies": {
    "openclaw": "*"
  }
}

Unverifiable Dependency: openclaw has 16 known advisory(ies) (CVE-2026-53846 (OpenClaw: Workspace .env npm_execpath could influence bundled runtime dependency); CVE-2026-32064 (OpenClaw's andbox browser noVNC observer lacked VNC authentication); CVE-2026-32006 (OpenClaw has a BlueBubbles group allowlist mismatch via DM pairing-store fallbac) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
90% confidence
Finding

The manifest references a dependency with known advisories but does not pin or constrain the version, making it impossible to verify whether deployments are exposed to those issues. Because this skill operates inside the OpenClaw environment and writes to the workspace, inherited platform vulnerabilities could directly affect confidentiality or integrity of user data.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
72% confidence
Finding

The comment emphasizes that text input is safely passed via stdin, which is true, but the surrounding documentation for the skill presents it as writing within the OpenClaw workspace while the code actually permits arbitrary output destinations. In context, this documentation contributes to a misleading impression of constrained, safe side effects that the implementation does not enforce.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.