Back to skill

Security audit

Voice TTS

Security checks for vulnerabilities and agentic risk

Overview

This skill is a straightforward Chinese text-to-speech helper, but users should understand it installs an external dependency, sends text to edge-tts, and writes generated audio files locally.

Install only if you are comfortable using edge-tts and sending the text you synthesize to that service. Prefer pipx with a reviewed pinned version, use the default ~/.openclaw/media output location, avoid secrets or regulated data in TTS text, and verify message destinations before sending generated audio.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:24
Finding

Unpinned Third-Party Dependency Installation

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/tts.sh:54
Finding

Caller-Controlled Output Path Allows Arbitrary File Overwrite

Content
View full analysis
/dev/null | grep -q "$VOICE"; then log_warn "声音 '$VOICE' 可能不存在,使用默认声音: zh-CN-YunxiNeural" VOICE="zh-CN-YunxiNeural" fi log_info "开始生成语音..." log_info " 声音: $VOICE" log_info " 文本: ${TEXT:0:50}${TEXT_LENGTH:50:+...}" log_info " 输出: $OUTPUT" # 生成语音(保留错误输出用于调试) if ! edge-tts \ --voice "$VOICE" \ --text "$TEXT" \ --write-media "$OUTPUT" 2>&1; then ``` ### Technical Analysis The third positional argument is accepted as an unrestricted output path and passed directly to `edge-tts --write-media`. The script does not canonicalize the path, require it to remain under `$HOME/.openclaw/media`, reject symbolic links, or prevent replacement of an existing file. The writability check applies only to the default `$OUTPUT_DIR`, even when the caller supplies a path in another directory. Consequently, any caller that can influence the third argument can direct audio output to any file writable by the process. A pre-positioned symbolic link can similarly redirect output to another writable target. The default filename uses `date +%s`, providing only one-second resolution. Two concurrent invocations within the same second can select the same output path, allowing one generated message to overwrite or corrupt another. Shell command injection is not present here because `"$OUTPUT"` is quoted and is passed as one argument. The issue is file- ...[truncated 1731 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (7)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

代码的核心功能是本地 TTS 文件生成,这与描述中的“使用 edge-tts 生成”部分一致,也确实支持选择不同 voice。但描述把能力扩展到了“发送语音消息到多渠道”,而代码完全没有网络发送、平台集成、Webhook/API 调用等逻辑;此外描述提到“可调节语速音调”,代码只处理文本、voice 和输出路径三个参数,没有任何速率或音调配置。因此描述高于实际实现,存在实质性能力不匹配。

Content

No source excerpt is available for this finding.

YARA rule 'backdoor_persistence': Backdoor persistence with malicious payloads (shell commands, SSH key injection, hidden root users) [malware]

High
Category
YARA Match
Confidence
75% confidence
Finding

YARA rule matched a known malware signature (reverse shell, backdoor, ransomware, C2 framework, or info stealer).

Content

Scanner excerpt · SKILL.md (reported line 46)May include surrounding context.

️ macOS 的 Homebrew Python 受 PEP 668 限制,禁止直接 pip install。

不要用 pip install --break-system-packages,会污染系统 Python。

Linux(推荐 pipx,或用 --user 安装):

bash
# 方式一:pipx(推荐)
pipx install edge-tts

# 方式二:pip --user(如果 pipx 不可用)
pip install --user edge-tts
# 安装后确保 ~/.local/bin 在 PATH 中
echo 'export PATH="$HOME/.local/bin:$PATH"' >> ~/.bashrc && source ~/.bashrc

Windows(PowerShell):

powershell
pip install edge-tts
# 或
pipx install edge-tts

3. 验证安装

bash
edge-tts --list-voices | head -5

看到语音列表即安装成功。


快速流程

  1. 运行脚本生成语音 → 获取文件路径
  2. 用 message 工具发送(asVoice: true,filePath)

生成语音

bash
bash <skill_dir>/scripts/tts.sh "文本内容" [voice] [output_path]

预期输出:

  • 返回生成的音频文件路径(如 `~/.openclaw/m

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The description says to use the skill whenever the user asks to '发语音、语音回复、TTS、文字转语音、语音播报、语音消息'. Several of these phrases are broad everyday requests, and the file does not provide scope limits, exclusions, or negative examples to clarify when this skill should or should not activate.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The skill does not clearly warn that user-provided text is sent to an external TTS provider, which can expose sensitive content to a third party. In a messaging skill, users may submit confidential notifications, internal announcements, or personal data, so the lack of disclosure meaningfully increases privacy risk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The usage text, examples, and default voice are all hard-coded for Chinese text and zh-CN voices, which imposes a specific language/locale by default. There is no natural-language indication that other languages are supported, no user choice prompt, and no documented justification that this skill is intentionally region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

The description repeatedly frames the skill as generating '中文语音' and lists only Chinese voices, but it does not explain that the skill is intentionally limited to Chinese or offer the user a language choice. This can be a language/locale policy issue when a skill forces a specific language without explicit opt-in or justification.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

前文参数说明在 L083-L086 明确写明第 1 参数是文本、第 2 参数是声音,但 L157 和 L173 的示例却把 voice 放在第 1 参数、文本放在第 2 参数。这会直接误导调用者,属于文档对实际接口意图的主动性矛盾,而非单纯信息缺失。

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.