Back to skill

Security audit

Qwen3-TTS Voice Synthesis

Security checks for vulnerabilities and agentic risk

Overview

This TTS skill is mostly coherent, but it can send speech text outside the local machine through automatic fallback or a configurable endpoint without strong user-facing controls.

Review before installing if you may synthesize private, proprietary, or user-sensitive text. Use it only with a trusted local ComfyUI endpoint, avoid untrusted COMFYUI_URL values, disable fallback with --fallback-edge false when text must stay local, and be aware that the advertised multi-role dialogue helper is not included in the inspected package.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/qwen_tts.py:20
Finding

Environment-Controlled ComfyUI Endpoint Enables SSRF and Speech-Content Disclosure

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/qwen_tts.py:217
Finding

Unvalidated ComfyUI Output Paths Permit Directory Traversal and Arbitrary Local Media Processing

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
Findings (17)

Tainted flow: 'req' from os.environ.get (line 179, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

The request target is derived from COMFYUI_URL in the environment and then used for an outbound HTTP request without validation. An attacker who can influence the environment can redirect requests to arbitrary internal or external endpoints, creating SSRF-style behavior and potentially exfiltrating synthesized text or probing services reachable from the host.

Content

Scanner excerpt · scripts/qwen_tts.py (reported line 185)May include surrounding context.

python
headers={"Content-Type": "application/json"},
    )
    try:
        with urllib.request.urlopen(req, timeout=30) as resp:
            result = json.loads(resp.read())
            return result.get("prompt_id")
    except (urllib.error.URLError, urllib.error.HTTPError) as e:

Tainted flow: 'url' from os.environ.get (line 198, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

This polling request uses the same environment-controlled COMFYUI_URL to build the history endpoint and fetches it without any trust boundary checks. In the skill context, the script is advertised as local TTS, so silent redirection to a remote service is more dangerous because users may assume text stays on-device while it can instead be transmitted elsewhere.

Content

Scanner excerpt · scripts/qwen_tts.py (reported line 199)May include surrounding context.

python
while time.time() - start < timeout:
        try:
            url = f"{COMFYUI_URL}/history/{prompt_id}"
            with urllib.request.urlopen(url, timeout=10) as resp:
                history = json.loads(resp.read())
                if prompt_id in history:
                    status = history[prompt_id].get("status", {})

Tp4

High
Category
MCP Tool Poisoning
Confidence
90% confidence
Finding

The code generally aligns with a local TTS purpose and does implement voice cloning and voice design, so the high-level domain is correct. However, there are material description-behavior mismatches. Most importantly, the script header and logic make clear it is '单角色语音合成' (single-role synthesis), not multi-character dialogue. It only accepts one --voice argument and produces one output audio file per invocation, with no dialogue orchestration, speaker switching, or role parsing. Also, the declared primary path mentions '琪琪OPC首选TTS', but the code's actual primary engine is ComfyUI-driven Qwen3-TTS; tts-cosyvoice/Edge TTS is used only as a fallback subprocess. Finally, while the description may suggest richer end-user support, the code explicitly states Qwen3-TTS does not directly support SRT generation and merely suggests separate handling. These are substantive enough to flag as a mismatch, though not a completely unrelated skill.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill advertises automatic fallback from local TTS to Edge TTS without clearly warning that user-provided text and possibly sensitive dialogue content may be sent to a cloud service. In a TTS skill, inputs often contain private conversations, names, or proprietary scripts, so silent off-device transmission creates a meaningful confidentiality risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The fallback strategy section describes automatic cloud TTS use but omits a clear notice about data leaving the local environment. Because this skill is positioned as a local Qwen3-TTS solution, users may reasonably expect all synthesis to remain on-device, making undocumented cloud fallback especially deceptive and privacy-impacting.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The skill is described as local TTS, but on failure it falls back to another skill described as Edge TTS-based, which may use a non-local service. This creates a privacy and trust-boundary violation because user text may leave the machine unexpectedly when local synthesis fails.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This code hard-codes Chinese locale behavior in two places: the fallback TTS voices are all zh-CN variants, and the CLI default for --language is "Chinese". That creates a natural-language locale policy concern because the skill effectively forces a specific language/locale unless the user notices and overrides it.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
74% confidence
Finding

In addition to submitting TTS jobs, the code invokes a separate Python script for fallback synthesis and runs ffmpeg as a subprocess for conversion. While audio conversion may be implementation-related, arbitrary subprocess execution is a stronger capability than the manifest's simple TTS description makes explicit.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/qwen_tts.py (reported line 249)May include surrounding context.

python
cmd.extend(["--srt", srt_path])

    import subprocess
    result = subprocess.run(cmd, capture_output=True, text=True)
    if result.returncode != 0:
        print(f"❌ Edge TTS 也失败了: {result.stderr}")
        return False

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/qwen_tts.py (reported line 261)May include surrounding context.

python
cmd.extend(["--srt", srt_path])

    import subprocess
    result = subprocess.run(cmd, capture_output=True, text=True)
    if result.returncode != 0:
        print(f"❌ Edge TTS 也失败了: {result.stderr}")
        return False

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
76% confidence
Finding

This JSON preset defines user-facing voice metadata primarily in Chinese, but the synthesis instruction at L16 is fixed in English. That can impose a language/locale requirement on downstream use without any documented choice or justification, which fits the policy category for forced language behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
74% confidence
Finding

The instruct field for this preset is fixed in English even though the surrounding preset metadata is Chinese. This may force a specific language or locale behavior in the skill configuration without opt-in or documented regional justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
74% confidence
Finding

The instruct field for this preset is fixed in English even though the surrounding preset metadata is Chinese. This may force a specific language or locale behavior in the skill configuration without opt-in or documented regional justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
74% confidence
Finding

The instruct field for this preset is fixed in English even though the surrounding preset metadata is Chinese. This may force a specific language or locale behavior in the skill configuration without opt-in or documented regional justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
74% confidence
Finding

The instruct field for this preset is fixed in English even though the surrounding preset metadata is Chinese. This may force a specific language or locale behavior in the skill configuration without opt-in or documented regional justification.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The module documentation explicitly states '单角色语音合成脚本', and the CLI accepts only one --voice value and synthesizes a single text stream per run. This does not match the broader advertised capability of '多角色对话' at the skill level, creating an intent mismatch between documentation and apparent scope.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.