Back to skill

Security audit

Text to Voice Local

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent text-to-speech wrapper, but it presents itself as local while using an Edge TTS dependency that may send text to an external service without clear disclosure.

Review before installing if you may convert private, confidential, or regulated text. Treat generated speech input as potentially sent to an external Edge TTS service, and prefer a pinned local dependency or an offline TTS engine if privacy or reproducibility matters.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
scripts/edge_tts.js:4
Finding

Unpinned Globally Loaded Third-Party TTS Dependency

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:76-79; scripts/edge_tts.js:4-8
Vulnerability Type: Unpinned third-party dependency loaded from global or fixed system locations
Risk Level: Medium

Complete Code Snippet

From SKILL.md:76-79:

bash
If `node-edge-tts` is missing:
```bash
npm i -g node-edge-tts
text

From `scripts/edge_tts.js:4-8`:

```javascript
const candidates = [
  'node-edge-tts',
  '/usr/lib/node_modules/openclaw/node_modules/node-edge-tts',
];

Technical Analysis

The documented installation command retrieves the latest available version of node-edge-tts without an exact version constraint, lockfile, or separately verified integrity value. The runtime then attempts to load either a module resolved by Node.js or a package from a fixed OpenClaw system path.

Because require() executes package initialization code, the effective trusted codebase includes whichever dependency copy resolves at runtime. The project does not define an audited local dependency version or provide a lockfile that makes installations reproducible. Consequently, a compromised future package release, compromised package-distribution account, or attacker-controlled package copy in an applicable resolution location could introduce arbitrary JavaScript execution.

This finding concerns supply-chain trust and reproducibility. The audited repository itself does not contain evidence that the current node-edge-tts package is malicious.

Attack Path

  1. An attacker compromises the upstream npm package, its publisher account, or a package copy available through an applicable runtime resolution location.
  2. A user follows the documented npm i -g node-edge-tts command, which does not select a previously audited exact version, or the environment otherwise contains a substituted copy.
  3. The user invokes the text-to-speech workflow.
  4. scripts/edge_tts.js calls require() on the dep ...[truncated 815 chars]
Remediation
View remediation

Remediation Suggestions

  1. Add a local package.json that pins node-edge-tts to an exact, reviewed version rather than using a floating global installation.
  2. Commit a generated lockfile and install dependencies with npm ci to ensure reproducible resolution.
  3. Review and retain the lockfile integrity metadata; use package provenance or registry-signature verification where supported.
  4. Load only the project-local dependency and remove global and absolute-path fallback resolution.
  5. Run the TTS process under a least-privileged account or sandbox with access limited to the required input and output locations.
  6. Incorporate dependency vulnerability, provenance, and unexpected lifecycle-script checks into release review.
  7. Document whether the dependency sends text to a remote service so users can make informed decisions before processing sensitive content.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (18)

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The skill is presented as local text-to-voice generation, but the documented dependency on node-edge-tts strongly suggests text may be sent to a network-backed TTS service instead of being processed fully offline. That mismatch can cause users to expose sensitive text under false assumptions about locality/privacy, and the inaccurate claims about canonical output and workflow behavior reduce operator trust and auditability.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 64)May include surrounding context.

md
- `scripts/edge_tts.js`

Chaining Abuse

High
Category
Tool Misuse
Confidence
75% confidence
Finding

Tool calls are chained to bypass individual safety checks or escalate capabilities beyond what any single tool call would allow.

Content

Scanner excerpt · scripts/tts_from_file_chunked.sh (reported line 21)May include surrounding context.

sh
tmp_ext="${tmp_base##*.}"
[ "$tmp_ext" = "$tmp_base" ] && tmp_ext="mp3"
TMP_FINAL="$WORKDIR/${tmp_stem}.concat.$$.$tmp_ext"
cleanup(){ rm -rf "$WORKDIR"; rm -f "$TMP_FINAL"; }
trap cleanup EXIT
python3 - "$INPUT_TXT" "$PARTS_TXT" <<'PY'
import sys, re, pathlib

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill is described as a local text-to-voice workflow, but it instantiates node-edge-tts, which relies on Microsoft's Edge TTS service rather than offline synthesis. This mismatch creates a security and privacy issue because users may provide sensitive text under the assumption it stays local, when it is actually sent off-host.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The code introduces an unnecessary external network dependency for a skill whose declared purpose is local-only voice generation. In an agent/workspace setting, this expands the trust boundary and can expose confidential text to third-party processing, making the behavior more dangerous than in a clearly cloud-based TTS skill.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The call to tts.ttsPromise(text, out) sends the provided text to an external TTS service with no warning, consent flow, or disclosure in the script. This is dangerous because arbitrary user or workspace text may include secrets, personal data, or internal content that is silently transmitted to a third party.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · scripts/text_to_voice.sh (reported line 74)May include surrounding context.

sh
if [ "$missing" -ne 0 ]; then
    printf '\nSuggested install commands:\n'
    [ "$node_status" = "missing" ] && printf -- '- install node via your system package manager or nvm\n'
    [ "$ffmpeg_status" = "missing" ] && printf -- '- sudo apt install ffmpeg\n'
    if [ "$edge_tts_status" = "missing" ]; then
      printf -- '- npm i -g node-edge-tts\n'
    fi

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script emits natural-language strings such as 'Озвучка текста', 'Старт', and 'Готово' directly to users, which imposes a specific locale without opt-in. The policy allows locale constraints only when documented and justified or when the user is given a choice, neither of which is present here.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This shell script performs a file write by moving the temporary output over the user-supplied destination path, which can overwrite an existing file. Aside from a usage line, there is no confirmation prompt, warning comment, or user-facing disclosure that the target file will be replaced.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script defaults VOICE to ru-RU-DmitryNeural, which imposes a specific language/locale choice when the caller does not supply one. The file does not offer opt-in language selection or explain why a Russian locale is required, which can violate language/locale policy expectations.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

This code creates temporary directories and files, invokes another script and ffmpeg as subprocesses, deletes temporary data with rm -rf, and overwrites the requested output path with mv -f. Within this shell file there is no confirmation prompt, user-facing warning, or explanatory comment/docstring disclosing these safety-relevant behaviors.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The default voice is set to ru-RU-DmitryNeural, which imposes a specific language/locale choice when the caller does not provide a voice argument. This is a natural-language policy concern because the script does not offer explicit user opt-in or explain why a Russian voice is required.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The script defaults VOICE to "ru-RU-DmitryNeural", which imposes a specific language/locale when the user does not provide a choice. This is a natural-language policy concern because the file does not indicate that the Russian locale is optional or justified for a region-specific tool.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The default_voice is hard-coded to ru-RU-DmitryNeural, which imposes a specific language/locale in the skill configuration. Forcing a locale without offering user choice or documenting a justified region-specific constraint matches the policy-violation criteria for language or locale restrictions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

The script defaults to the specific locale-bound voice 'ru-RU-DmitryNeural', which imposes a language/locale choice unless the user overrides it manually. There is no surrounding natural-language justification or opt-in mechanism indicating that Russian output is required or intended by policy.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

This code creates directories under the workspace and initializes a persistent state file, and later the skill is documented to produce a canonical MP3 output. While these writes are part of the tool's function, the script itself provides no explicit warning or disclosure at the point of operation about creating persistent files in the workspace.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

The script defaults to ru-RU-DmitryNeural, which imposes a specific language/locale choice when the caller does not provide a voice. That is a natural-language locale constraint without any opt-in, alternative selection flow, or documented region-specific justification in this file.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.