Back to skill

Security audit

Feishu Voice Message

Security checks for vulnerabilities and agentic risk

Overview

This skill has a coherent audio-generation purpose, but it overstates Feishu sending and runs mutable, unpinned external code from another skill path.

Review this before installing. Treat it as a local audio-file generator, not a verified Feishu sender. Use only non-sensitive text unless you have checked the Edge TTS data flow, pin or inspect the external converter it executes, and keep outputs in a dedicated directory to avoid accidental overwrites.

Vulnerability Patterns
  • Tool Hijacking and SpoofingModifies or replaces tools so legitimate-looking calls execute attacker logic
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T07 · Tool Hijacking and Spoofing

Error
Location
scripts/feishu_voice.py:41
Finding

Execution of an Unpinned and Externally Mutable TTS Component

Content
View full analysis

Vulnerability Details

File Location: scripts/feishu_voice.py:41-50; related dependency declarations at skill-info.json:22-24 and SKILL.md:83-86
Vulnerability Type: Untrusted dependency execution and local tool substitution
Risk Level: High

Vulnerable Code

python
cmd = [
    "node",
    os.path.expanduser("~/.openclaw/workspace/skills/edge-tts/scripts/tts-converter.js"),
    text,
    "--voice", config["voice"],
    "--pitch", config["pitch"],
    "--rate", config["rate"],
    "--output", output_mp3
]
python
result = subprocess.run(cmd, capture_output=True, text=True, timeout=30)

The dependency is declared without an exact version:

json
"dependencies": {
  "edge-tts": ">=0.1.0"
}

The documented installation command is also unpinned:

bash
npm install edge-tts

Technical Analysis

The Python script executes tts-converter.js from a fixed path under another skill directory in the current user's workspace. That JavaScript file is not included in this project and therefore was not available for inspection as part of this audit.

The dependency declaration permits any version greater than or equal to 0.1.0, and the documented installation command does not use a lockfile, exact version, or integrity verification. In addition, the code does not invoke a verified package entry point. It directly trusts a file in an independently writable local directory.

Consequently, the effective executable payload can differ from the code reviewed here. Any process or user capable of replacing that JavaScript file, changing the target through a filesystem link, or influencing the installed dependency can cause arbitrary JavaScript to be executed when the skill is invoked.

Although subprocess arguments are passed as an array and do not create direct shell injection, this does not protect against substitution of the executable JavaScript file itself.

...[truncated 1320 chars]

Remediation
View remediation

Remediation Suggestions

  1. Bundle the required converter with this skill and include it in security review, or invoke a package through its documented and verified package entry point.
  2. Pin the dependency to an exact reviewed version rather than accepting >=0.1.0.
  3. Commit and enforce an appropriate lockfile containing integrity hashes.
  4. Install dependencies using a reproducible, integrity-verifying command such as a frozen-lockfile installation.
  5. Before execution, resolve the converter path and reject symbolic links, unexpected file ownership, or paths outside an approved installation directory.
  6. Verify the converter against a trusted cryptographic digest before each execution or deploy it in a read-only package directory.
  7. Avoid relying on another independently mutable skill directory for executable code.
  8. Run the converter with least privilege, a restricted environment, and limited filesystem and network access.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/feishu_voice.py:94
Finding

Arbitrary Writable Audio-File Overwrite Through User-Controlled Output Paths

Content
View full analysis

Vulnerability Details

File Location: scripts/feishu_voice.py:72-78 and scripts/feishu_voice.py:94-111
Vulnerability Type: Unrestricted output path and unsafe file overwrite
Risk Level: Medium

Vulnerable Code

python
cmd = [
    "ffmpeg",
    "-i", input_mp3,
    "-c:a", "libopus",
    "-b:a", "64k",
    "-y",
    output_opus
]
python
parser.add_argument('--output', '-o', default=None, help='输出路径(可选)')

args = parser.parse_args()

if args.output:
    base_path = args.output.replace('.opus', '').replace('.mp3', '')
else:
    base_path = os.path.join(os.environ.get('TEMP', '/tmp'), 'openclaw', 'voice_message')

os.makedirs(os.path.dirname(base_path) if os.path.dirname(base_path) else '.', exist_ok=True)

mp3_path = f"{base_path}.mp3"
opus_path = f"{base_path}.opus"

if not generate_tts(args.text, mp3_path, args.preset):
    sys.exit(1)

if not convert_to_opus(mp3_path, opus_path):
    sys.exit(1)

Technical Analysis

The --output argument is accepted as an unrestricted filesystem path. The code does not canonicalize the destination, constrain it to a dedicated output directory, inspect existing files, or reject symbolic links.

Two files are derived from the attacker-selected base path: an MP3 destination passed to the external TTS converter and an OPUS destination passed to FFmpeg. FFmpeg receives -y, which explicitly permits overwriting an existing OPUS file without confirmation. The external TTS converter may similarly replace the MP3 destination, although its implementation was not included and its precise overwrite behavior could not be verified.

The extension handling uses unrestricted string replacement:

python
args.output.replace('.opus', '').replace('.mp3', '')

This removes matching text anywhere in the supplied path rather than safely replacing only the final suffix. It can therefore produce an unexpected destination path.

...[truncated 1677 chars]

Remediation
View remediation

Remediation Suggestions

  1. Create a dedicated output directory with restrictive permissions and require every generated file to remain under that directory.
  2. Resolve the requested path with pathlib.Path.resolve() and verify it is a descendant of the approved output directory.
  3. Use Path.with_suffix() to handle the final extension rather than unrestricted string replacement.
  4. Reject symbolic links for the output directory, intermediate MP3 file, and final OPUS file.
  5. Refuse to overwrite existing files by default. Remove FFmpeg's -y option or replace it with -n.
  6. If overwriting is a required feature, require an explicit trusted confirmation flag and verify the destination immediately before use.
  7. Generate unpredictable temporary filenames using tempfile.TemporaryDirectory or NamedTemporaryFile.
  8. Where supported, create destinations atomically with exclusive-create semantics to reduce race conditions.
  9. Remove the intermediate MP3 securely after successful conversion and on failure.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (9)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill advertises direct Feishu voice-message generation and sending, but the described behavior appears limited to local file production plus manual follow-up. This mismatch can mislead users about what data is transmitted, what automation occurs, and what trust boundary applies, which is dangerous in agent systems because users may approve actions under false assumptions.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill references shell execution and environment-dependent tooling (npm, ffmpeg, python) but does not declare any explicit tool scope such as allowed-tools or permissions. In an agent environment, this weakens least-privilege controls and can let a seemingly simple content-generation skill invoke broader local execution capabilities than users expect.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The trigger list contains broad, common phrases like TTS, voice message, and 发送语音 that may match ordinary conversation. Overbroad activation can cause unintended invocation of shell-backed media generation workflows, increasing the chance of accidental execution, surprise file creation, or unreviewed processing of sensitive text.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SKILL.md (reported line 80)May include surrounding context.

bash
   # Windows: Download from https://ffmpeg.org
   # Mac: brew install ffmpeg
   # Linux: sudo apt install ffmpeg

Technical Details / 技术细节

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

The skill discusses sending generated audio to Feishu but does not clearly warn users that the synthesized content may be uploaded or transmitted to a third-party platform. This omission creates a privacy and consent risk, especially if users provide sensitive text assuming processing remains local.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The module documentation explicitly states '自动发送到对话' as a listed capability. However, the implementation only performs TTS generation and OPUS conversion, then prints a note telling the user to manually send the .opus file to Feishu; there is no Feishu API call or chat-delivery logic anywhere in the file.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This code hard-codes only zh-CN voice presets and presents all user-facing descriptions in Chinese, which creates a locale-specific constraint without user opt-in. Under the policy rule, forcing a specific language or locale is a natural-language policy violation unless the tool explicitly offers a choice or clearly documents a justified region-specific limitation.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/feishu_voice.py (reported line 55)May include surrounding context.

python
print(f"   预设: {preset} ({config['voice']})")
    
    try:
        result = subprocess.run(cmd, capture_output=True, text=True, timeout=30)
        if result.returncode == 0:
            print(f"✅ MP3 已生成: {output_mp3}")
            return True

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/feishu_voice.py (reported line 80)May include surrounding context.

python
print(f"   预设: {preset} ({config['voice']})")
    
    try:
        result = subprocess.run(cmd, capture_output=True, text=True, timeout=30)
        if result.returncode == 0:
            print(f"✅ MP3 已生成: {output_mp3}")
            return True

Static analysis

No suspicious patterns detected.