Back to skill

Security audit

Edge Tts Zh

Security checks for vulnerabilities and agentic risk

Overview

This skill mostly implements Chinese text-to-speech, but it also ships undisclosed helper scripts that can rewrite an installed skill and inject hidden PowerShell playback behavior.

Review this package before installing. The basic TTS function is recognizable, but remove or ignore patch.py, fix.py, and speak_temp.txt unless you explicitly want them, and avoid running those helpers because they can rewrite an installed skill and introduce hidden PowerShell execution. Do not synthesize secrets or sensitive text unless you are comfortable sending it to the external Edge TTS service, and expect generated audio to play automatically if pygame is installed.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (4)

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
scripts/fix.py:2
Finding

Auxiliary scripts overwrite an installed Skill outside the project directory

Content
View full analysis

Vulnerability Details

File Location: scripts/fix.py:2-13; scripts/patch.py:5-28
Vulnerability Type: Arbitrary modification of an installed OpenClaw Skill
Risk Level: Medium

Vulnerable Code

scripts/fix.py:2-3,12-13:

python
with open('C:/Users/Administrator/.openclaw/workspace/skills/edge-tts-zh/scripts/speak.py', 'r', encoding='utf-8') as f:
    lines = f.readlines()

with open('C:/Users/Administrator/.openclaw/workspace/skills/edge-tts-zh/scripts/speak.py', 'w', encoding='utf-8') as f:
    f.writelines(lines)

scripts/patch.py:5-6,24-28:

python
with open('C:/Users/Administrator/.openclaw/workspace/skills/edge-tts-zh/scripts/speak.py', 'r', encoding='utf-8') as f:
    content = f.read()

content = content.replace(old_code, new_code)

with open('C:/Users/Administrator/.openclaw/workspace/skills/edge-tts-zh/scripts/speak.py', 'w', encoding='utf-8') as f:
    f.write(content)

Technical Analysis

Both scripts use an absolute path to read and overwrite a separate installed copy of speak.py in the OpenClaw workspace. The target is not resolved relative to the current package, and the operation does not request confirmation, validate the target, create a backup, or use an atomic replacement.

This crosses the expected boundary of a text-to-speech Skill. Running either development helper can alter executable code used by future invocations of the installed Skill. The operation executes with the current process's filesystem privileges; it does not independently elevate operating-system privileges.

Attack Path

  1. The user or Agent invokes scripts/fix.py or scripts/patch.py.
  2. The script accesses the hardcoded installed-Skill location under the Administrator profile.
  3. It modifies the installed speak.py rather than a file scoped to the current project.
  4. Future calls to the installed Skill execute the modified implementation.
  5. If the replacement is in ...[truncated 499 chars]
Remediation
View remediation

Remediation Suggestions

  • Remove fix.py and patch.py from the distributed package if they are development-only utilities.
  • Resolve modification targets relative to the current package rather than using an absolute user-profile path.
  • If external modification is necessary, require an explicit target argument and interactive confirmation.
  • Canonicalize the target with Path.resolve() and enforce that it is within an approved directory.
  • Create a backup and use an atomic temporary-file replacement.
  • Verify that the expected source content is present and abort if the installed version is not compatible.
  • Document every file that the utility may modify.

T09 · Insecure Skill Coding Practices

Error
Location
scripts/patch.py:9
Finding

Patch helpers introduce PowerShell command injection through the output path

Content
View full analysis

Vulnerability Details

File Location: scripts/fix.py:7-9; scripts/patch.py:9-24; scripts/speak_temp.txt:4-16
Vulnerability Type: PowerShell command injection
Risk Level: High

Vulnerable Code

scripts/patch.py:9-24:

python
old_code = '        print(output_path)  # stdout 输出文件路径\n        return True'
new_code = '''        # 自动播放(后台静默播放,不弹窗)
        try:
            subprocess.run(
                ["powershell", "-Command", f'Start-Process "{output_path}" -WindowStyle Hidden'],
                capture_output=True,
                timeout=5
            )
            print(f"🔊 正在播放...", file=sys.stderr)
        except Exception as e:
            print(f"⚠️  播放失败:{e}", file=sys.stderr)

        print(output_path)  # stdout 输出文件路径
        return True'''

content = content.replace(old_code, new_code)

scripts/fix.py:7-9:

python
if 'subprocess.run' in line and 'Start-Process' in line:
    lines[i] = '        subprocess.run(["powershell", "-Command", f\'Start-Process "{output_path}" -WindowStyle Hidden\'], capture_output=True, timeout=5)\n'
    break

Technical Analysis

The helper scripts inject a PowerShell command in which output_path is interpolated directly into PowerShell source code:

python
f'Start-Process "{output_path}" -WindowStyle Hidden'

In the patched program, output_path originates from the user-controlled --output option. Passing PowerShell as an argument array does not prevent injection because the value following -Command is still parsed as PowerShell source.

PowerShell performs expression expansion inside double-quoted strings, including subexpressions using $(). Characters used for PowerShell subexpressions can be present in a Windows filename. Consequently, a crafted output path can remain acceptable as a filesystem path while causing PowerShell to evaluate attacker-controlled expressions when playback begins.

The ...[truncated 1403 chars]

Remediation
View remediation

Remediation Suggestions

  • Remove the PowerShell-based playback patch and the obsolete helper artifacts.
  • Never interpolate a user-controlled path into the source passed to powershell -Command.
  • Prefer a native Python playback API with no shell interpreter.
  • If PowerShell is unavoidable, pass the path as separately encoded data to a fixed script and use strict literal handling rather than source interpolation.
  • Canonicalize output paths and reject unexpected control characters or shell-expression syntax as defense in depth.
  • Restrict output to a designated directory and generate server-controlled filenames.
  • Add security tests using paths containing $(), quotes, semicolons, pipes, and newline characters.

T08 · Insecure Dependencies

Note
Location
SKILL.md:63
Finding

Third-party dependency installation is unpinned and unverifiable

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:63-67; README.md:16; scripts/speak.py:40-44
Vulnerability Type: Unpinned third-party dependency
Risk Level: Low

Vulnerable Code

SKILL.md:63-67:

bash
pip install edge-tts

scripts/speak.py:40-44:

python
try:
    import edge_tts
except ImportError:
    print("❌ edge-tts 未安装,运行:pip install edge-tts", file=sys.stderr)
    return False

Technical Analysis

The installation guidance retrieves the mutable latest release of edge-tts. The project does not provide a pinned version, dependency lockfile, package hash, or explicit trusted package index.

This makes installation non-reproducible. A future compromised, malicious, or incompatible release could be installed without any project change or additional review. Python package installation may also execute package-controlled build logic, while imported dependency code executes with the same privileges as this Skill.

No evidence was found that the currently named edge-tts package is malicious. The finding concerns the unsafe and unverifiable dependency-resolution process.

Attack Path

  1. A user follows the documented pip install edge-tts command.
  2. pip resolves whichever package version is current at installation time.
  3. If the package source or a future release is compromised, malicious package-controlled code is installed.
  4. Build-time logic may execute during installation.
  5. Dependency code subsequently executes when speak.py imports and invokes edge_tts.

Impact Assessment

A compromised dependency can execute code with the installing or invoking user's privileges. This may expose files, environment variables, network access, and other resources available to the Python process. The practical likelihood is reduced because no malicious dependency or unsafe alternate package source was identified in the audited files.

Remediation
View remediation

Remediation Suggestions

  • Pin edge-tts to a reviewed version.
  • Add a requirements or lock file containing cryptographic hashes.
  • Install with hash enforcement, such as pip install --require-hashes.
  • Explicitly use a trusted package index and disable unneeded supplemental indexes.
  • Review dependency updates before changing the pin.
  • Document pygame as an optional dependency if playback remains supported.
  • Perform dependency vulnerability and provenance checks in continuous integration.

other

Note
Location
scripts/speak.py:49
Finding

Audio playback occurs automatically without a documented opt-in control

Content
View full analysis

Vulnerability Details

File Location: scripts/speak.py:49-59; scripts/speak.py:64-81
Vulnerability Type: Undisclosed automatic host-side effect
Risk Level: Low

Vulnerable Code

scripts/speak.py:49-59:

python
try:
    communicate = edge_tts.Communicate(text=text, voice=voice_name, rate=rate)
    await communicate.save(output_path)

    size = os.path.getsize(output_path)
    print(f"Generated: {output_path} ({size/1024:.1f}KB)", file=sys.stderr)

    # 自动播放(使用 pygame,完全后台无窗口)
    play_audio(output_path)

    return True

scripts/speak.py:64-81:

python
def play_audio(file_path):
    """使用 pygame 播放音频,完全隐藏窗口"""
    try:
        import pygame
        pygame.mixer.init()
        pygame.mixer.music.load(file_path)
        pygame.mixer.music.play()

        # 等待播放完成
        while pygame.mixer.music.get_busy():
            time.sleep(0.5)

        print("Playing completed!", file=sys.stderr)
        pygame.mixer.quit()
    except ImportError:
        print("⚠️ pygame not found, skipping auto-play", file=sys.stderr)
    except Exception as e:
        print(f"⚠️ Play failed: {e}", file=sys.stderr)

Technical Analysis

Every successful synthesis unconditionally calls play_audio(). There is no --play option, confirmation prompt, configuration control, or documented warning that ordinary synthesis will access the host audio device and block until playback finishes.

This behavior exceeds the documented generation and file-saving workflow. It is not arbitrary code execution, but it creates an unexpected host-side effect and can make unattended or automated invocation block for the duration of attacker-supplied text.

Attack Path

  1. A user or Agent invokes the documented text-to-speech command.
  2. The supplied text is synthesized and written to the output file.
  3. The script automatically initializes the host audio mixer.
  4. The generated ...[truncated 547 chars]
Remediation
View remediation

Remediation Suggestions

  • Disable playback by default.
  • Add an explicit --play option for users who want immediate playback.
  • Document the optional playback behavior and its pygame dependency.
  • Avoid waiting indefinitely; enforce a configurable maximum playback duration.
  • Ensure mixer resources are released with a finally block.
  • Allow noninteractive and server environments to disable all audio-device access.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (20)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description says the skill provides Chinese text-to-speech features such as speech synthesis, voice options, speed/pitch control, and subtitle generation. However, the provided code chunk does not synthesize speech, process text, generate subtitles, or expose TTS behavior. Instead, it edits another Python file on disk by replacing a subprocess invocation line. This is a materially different primary purpose: code patching/maintenance. While such a script could theoretically support the skill internally, this specific chunk performs undeclared file modification behavior and does not itself implement the described user-facing TTS capability.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description focuses on text-to-speech features: converting text to speech, selecting voices, adjusting speed/pitch, and generating subtitles. However, the provided code does not perform TTS synthesis or expose those features. Instead, it reads and rewrites a separate Python file at a hardcoded path, replacing code to add automatic background playback using PowerShell. File patching and process launching are materially different from the declared primary purpose and introduce undeclared capabilities. While audio playback could be related to a TTS workflow, this code’s actual function is to modify another script and trigger playback, which is not accurately represented by the declared description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The core purpose matches: the script uses Microsoft Edge TTS without an API key, supports multiple Chinese voices, and allows rate adjustment. However, there are notable description-behavior mismatches. First, the code automatically plays the generated audio using pygame, which is an undeclared capability beyond simple synthesis. Second, the description claims support for pitch adjustment and subtitle generation, but the code only exposes a rate parameter and contains no pitch handling or subtitle generation logic. Therefore the description does not accurately represent the actual behavior.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The description and title present the skill strictly as a Chinese text-to-speech tool (e.g., '中文语音合成') without any opt-in or explanation that the locale restriction is required for a region-specific use case. Under the policy, forcing a specific language without user choice or documented justification is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The script directly edits another file in the skill and specifically preserves a subprocess-based launch path using PowerShell Start-Process. In a TTS skill, silently patching executable behavior in a separate script is unnecessary for core text-to-speech functionality and increases the attack surface by retaining process-spawning behavior that could be abused or later modified for arbitrary execution.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The script overwrites another skill file in place with no prompt, backup, integrity check, or user consent. Silent code modification is dangerous because it hides behavioral changes, defeats review expectations, and can be used to persist unsafe functionality or tamper with trusted code paths.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This script modifies speak.py in place to add behavior not described in the skill metadata: automatic playback after synthesis. Self-modifying or patch-style behavior that silently changes another source file is risky because it obscures the effective behavior of the skill, bypasses normal review expectations, and can be repurposed to inject more harmful logic later.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The patch injects a hidden PowerShell Start-Process call to automatically play the generated audio with -WindowStyle Hidden, which introduces concealed process execution unrelated to the core requirement of generating TTS output. Hiding the spawned process and not disclosing the behavior increases user surprise and creates a precedent for silent execution of arbitrary files or commands if the output path or surrounding code is later influenced.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The injected snippet launches a subprocess through PowerShell in hidden-window mode and provides no disclosure to the user in this script. Hidden subprocess execution is dangerous because it reduces transparency, complicates detection, and can be abused to run additional commands beyond simple media playback.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script overwrites speak.py in place without backup, integrity checking, confirmation, or user notice. While not inherently code execution by itself, this is unsafe operationally because it silently alters installed behavior and can break the skill or conceal unauthorized modifications.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The module description and defaults explicitly position the skill as Chinese-only, including Chinese voices and a Chinese-language interface, but there is no opt-in or user-selectable locale behavior described. This can violate language/locale policy when a skill imposes a specific language by default without offering choice or documenting a justified regional constraint.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The manifest says the skill supports '字幕生成', but the implementation only calls edge_tts.Communicate(...).save(output_path) to write an MP3 and has no logic to generate subtitle/transcript files such as SRT or VTT. This is a direct mismatch between the advertised functionality and the actual behavior in this file.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest states that the skill supports both '语速/音调调节', but the CLI and synthesize call only accept and pass a rate parameter. There is no pitch argument in the parser and no pitch value supplied to the TTS engine, so the advertised pitch-control capability is not implemented here.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This file contains user-visible status and error messages such as "已生成", "正在播放", and "播放失败" that force a specific language. The policy requires avoiding language/locale constraints unless the user is given a choice or the limitation is explicitly justified.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The code goes beyond text-to-speech generation by automatically launching the produced file for playback. This creates an unexpected side effect on the host system and can be abused to trigger local process execution or disruptive behavior without explicit user consent, which is riskier than the stated TTS-only purpose suggests.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill invokes PowerShell and uses Start-Process to open the generated file, introducing an unnecessary process-spawning capability for a TTS function. Even though output_path appears intended as a local media file, launching a shell expands the attack surface and can become dangerous if the path is influenced unexpectedly or if host execution policies are permissive.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

The README promotes a cloud-backed TTS capability but does not disclose that user-provided text may be transmitted to Microsoft's Edge TTS service over the network. This can expose sensitive prompts, messages, or internal content to a third party, especially in agent workflows where users may assume processing is local because the skill emphasizes 'free' and 'no API key' rather than remote data handling.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The script performs automatic audio playback as a side effect after generating the file, which goes beyond a simple text-to-speech conversion utility. In agent or automation contexts, unexpected playback can leak synthesized sensitive content aloud, disrupt unattended workflows, or create an unanticipated side channel in shared environments.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The inline comment and docstring describe playback as '完全后台无窗口' / '完全隐藏窗口', which suggests a non-intrusive background action. In reality, the function loads and plays audio through pygame and then waits in a loop until playback finishes, producing an audible and blocking side effect that is materially different from the stated intent.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.