T09 · Insecure Skill Coding Practices
- Location
scripts/qwen_tts.py:20- Finding
Environment-Controlled ComfyUI Endpoint Enables SSRF and Speech-Content Disclosure
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This TTS skill is mostly coherent, but it can send speech text outside the local machine through automatic fallback or a configurable endpoint without strong user-facing controls.
Review before installing if you may synthesize private, proprietary, or user-sensitive text. Use it only with a trusted local ComfyUI endpoint, avoid untrusted COMFYUI_URL values, disable fallback with --fallback-edge false when text must stay local, and be aware that the advertised multi-role dialogue helper is not included in the inspected package.
scripts/qwen_tts.py:20Environment-Controlled ComfyUI Endpoint Enables SSRF and Speech-Content Disclosure
scripts/qwen_tts.py:217Unvalidated ComfyUI Output Paths Permit Directory Traversal and Arbitrary Local Media Processing
The request target is derived from COMFYUI_URL in the environment and then used for an outbound HTTP request without validation. An attacker who can influence the environment can redirect requests to arbitrary internal or external endpoints, creating SSRF-style behavior and potentially exfiltrating synthesized text or probing services reachable from the host.
headers={"Content-Type": "application/json"},
)
try:
with urllib.request.urlopen(req, timeout=30) as resp:
result = json.loads(resp.read())
return result.get("prompt_id")
except (urllib.error.URLError, urllib.error.HTTPError) as e:
This polling request uses the same environment-controlled COMFYUI_URL to build the history endpoint and fetches it without any trust boundary checks. In the skill context, the script is advertised as local TTS, so silent redirection to a remote service is more dangerous because users may assume text stays on-device while it can instead be transmitted elsewhere.
while time.time() - start < timeout:
try:
url = f"{COMFYUI_URL}/history/{prompt_id}"
with urllib.request.urlopen(url, timeout=10) as resp:
history = json.loads(resp.read())
if prompt_id in history:
status = history[prompt_id].get("status", {})
The code generally aligns with a local TTS purpose and does implement voice cloning and voice design, so the high-level domain is correct. However, there are material description-behavior mismatches. Most importantly, the script header and logic make clear it is '单角色语音合成' (single-role synthesis), not multi-character dialogue. It only accepts one --voice argument and produces one output audio file per invocation, with no dialogue orchestration, speaker switching, or role parsing. Also, the declared primary path mentions '琪琪OPC首选TTS', but the code's actual primary engine is ComfyUI-driven Qwen3-TTS; tts-cosyvoice/Edge TTS is used only as a fallback subprocess. Finally, while the description may suggest richer end-user support, the code explicitly states Qwen3-TTS does not directly support SRT generation and merely suggests separate handling. These are substantive enough to flag as a mismatch, though not a completely unrelated skill.
Without declared permissions the skill's intent is opaque and cannot be validated.
The skill advertises automatic fallback from local TTS to Edge TTS without clearly warning that user-provided text and possibly sensitive dialogue content may be sent to a cloud service. In a TTS skill, inputs often contain private conversations, names, or proprietary scripts, so silent off-device transmission creates a meaningful confidentiality risk.
The fallback strategy section describes automatic cloud TTS use but omits a clear notice about data leaving the local environment. Because this skill is positioned as a local Qwen3-TTS solution, users may reasonably expect all synthesis to remain on-device, making undocumented cloud fallback especially deceptive and privacy-impacting.
The skill is described as local TTS, but on failure it falls back to another skill described as Edge TTS-based, which may use a non-local service. This creates a privacy and trust-boundary violation because user text may leave the machine unexpectedly when local synthesis fails.
This code hard-codes Chinese locale behavior in two places: the fallback TTS voices are all zh-CN variants, and the CLI default for --language is "Chinese". That creates a natural-language locale policy concern because the skill effectively forces a specific language/locale unless the user notices and overrides it.
In addition to submitting TTS jobs, the code invokes a separate Python script for fallback synthesis and runs ffmpeg as a subprocess for conversion. While audio conversion may be implementation-related, arbitrary subprocess execution is a stronger capability than the manifest's simple TTS description makes explicit.
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
cmd.extend(["--srt", srt_path])
import subprocess
result = subprocess.run(cmd, capture_output=True, text=True)
if result.returncode != 0:
print(f"❌ Edge TTS 也失败了: {result.stderr}")
return False
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
cmd.extend(["--srt", srt_path])
import subprocess
result = subprocess.run(cmd, capture_output=True, text=True)
if result.returncode != 0:
print(f"❌ Edge TTS 也失败了: {result.stderr}")
return False
This JSON preset defines user-facing voice metadata primarily in Chinese, but the synthesis instruction at L16 is fixed in English. That can impose a language/locale requirement on downstream use without any documented choice or justification, which fits the policy category for forced language behavior.
The instruct field for this preset is fixed in English even though the surrounding preset metadata is Chinese. This may force a specific language or locale behavior in the skill configuration without opt-in or documented regional justification.
The instruct field for this preset is fixed in English even though the surrounding preset metadata is Chinese. This may force a specific language or locale behavior in the skill configuration without opt-in or documented regional justification.
The instruct field for this preset is fixed in English even though the surrounding preset metadata is Chinese. This may force a specific language or locale behavior in the skill configuration without opt-in or documented regional justification.
The instruct field for this preset is fixed in English even though the surrounding preset metadata is Chinese. This may force a specific language or locale behavior in the skill configuration without opt-in or documented regional justification.
The module documentation explicitly states '单角色语音合成脚本', and the CLI accepts only one --voice value and synthesizes a single text stream per run. This does not match the broader advertised capability of '多角色对话' at the skill level, creating an intent mismatch between documentation and apparent scope.
No suspicious patterns detected.