Back to skill

Security audit

三剪客 · 语音TTS

Security checks for vulnerabilities and agentic risk

Overview

This is a mostly disclosed voice/TTS integration, but its bundled client can use the user’s API key to call any a7w marketplace plugin, not just the documented voice endpoints.

Install only if you are comfortable giving this client a spend-capable a7w API key and understand that the bundled script can call any plugin available to that key, not only voice_tts. Use a limited or separate key if possible, monitor charges, avoid submitting sensitive audio or text unless you trust the provider’s data handling, and only clone voices with clear authorization from the speaker.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (28)

Tainted flow: 'req' from os.environ.get (line 151, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/client.py (reported line 111)May include surrounding context.

python
headers["Content-Type"] = "application/json"
    req = urllib.request.Request(url, data=data, headers=headers, method=method)
    try:
        with urllib.request.urlopen(req, timeout=timeout) as resp:
            raw = resp.read().decode("utf-8", "replace")
            status = resp.status
    except urllib.error.HTTPError as exc:

Tainted flow: 'req' from os.environ.get (line 151, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/client.py (reported line 153)May include surrounding context.

python
headers["Content-Type"] = "application/json"
    req = urllib.request.Request(url, data=data, headers=headers, method=method)
    try:
        with urllib.request.urlopen(req, timeout=timeout) as resp:
            raw = resp.read().decode("utf-8", "replace")
            status = resp.status
    except urllib.error.HTTPError as exc:

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared purpose is a voice/TTS skill, but the bundled client reportedly supports enumerating all apps, inspecting arbitrary schemas, calling arbitrary app/api pairs, exporting all plugin schemas, querying tasks/consumption, and storing API keys locally. That is a material expansion of capability beyond the stated purpose and could be abused as a general plugin-market proxy using the user's key, with access and data exposure users did not knowingly authorize.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 5)May include surrounding context.

md
多端点 AI 语音能力。支持 文字转语音(Live·异步)、克隆音色、文字转语音、文字转语音(异步)、语音转文字、音色列表。包内含完整操作文档与零依赖客户端(`SKILL.md` + `references/`)。需要自备 api.a7w.cn 的 API Key,注册领 Key 见 https://api.a7w.c

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 7)May include surrounding context.

md
多端点 AI 语音能力。支持 文字转语音(Live·异步)、克隆音色、文字转语音、文字转语音(异步)、语音转文字、音色列表。包内含完整操作文档与零依赖客户端(`SKILL.md` + `references/`)。需要自备 api.a7w.cn 的 API Key,注册领 Key 见 https://api.a7w.c

Missing User Warnings

High
Category
Not specified by scanner
Confidence
94% confidence
Finding

This documentation describes cloning a voice from uploaded audio or a remote audio URL without any warning about consent, privacy, or the biometric sensitivity of voice data. Voice cloning can enable impersonation, fraud, and non-consensual biometric processing, and the absence of safeguards in the skill documentation makes misuse materially more likely.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The call subcommand accepts arbitrary app and api names and forwards attacker-controlled JSON bodies to the remote marketplace, enabling the skill to perform any action exposed by the account rather than only voice-studio operations. In an agent context, this materially broadens the blast radius: prompt-induced misuse of the skill can trigger unrelated remote actions, data access, or spending against the user's account.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The entire skill description, usage instructions, safety notes, and contact information are presented only in Chinese, which constitutes a language/locale constraint in the skill's natural-language interface. The file does not offer an alternative language, opt-in choice, or any justification that this skill is intended only for a Chinese-speaking or region-specific audience.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill advertises executable client usage that can access environment, local files, network, and shell-like capabilities, but it does not declare any explicit tool scope or permissions boundaries. In an agent ecosystem, this makes the operational trust boundary unclear and increases the chance that the skill is run with broader privileges than users expect.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · README.md (reported line 11)May include surrounding context.

md
name: voice-tts-studio
slug: voice-tts-studio
displayName: 三剪客 · 语音TTS
description: "语音克隆、文字转语音(同步/异步)、语音识别等多端点 AI 语音能力。支持 文字转语音(Live·异步)、克隆音色、文字转语音、文字转语音(异步)、语音转文字、音色列表。包内含完整操作文档与零依赖客户端(`SKILL.md` + `references/`)。需要自备 api.a7w.cn 的 API Key,注册领 Key 见 https://api.a7w.cn/ 。遇到问题可加技术微信 9872659。"
version: 1.0.6
summary: "「语音TTS」的完整调用封装:6 个接口的官方文档、参数表与一个零依赖客户端。包内含完整操作文档与零依赖客户端(`SKILL.md` + `references/`)。需要自备 api.a7w.cn 的 API Key,注册领 Key 见 https://api.a7w.cn/ 。遇到问题可加技术微信 9872659。"
license: MIT

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · README.md (reported line 95)May include surrounding context.

md
name: voice-tts-studio
slug: voice-tts-studio
displayName: 三剪客 · 语音TTS
description: "语音克隆、文字转语音(同步/异步)、语音识别等多端点 AI 语音能力。支持 文字转语音(Live·异步)、克隆音色、文字转语音、文字转语音(异步)、语音转文字、音色列表。包内含完整操作文档与零依赖客户端(`SKILL.md` + `references/`)。需要自备 api.a7w.cn 的 API Key,注册领 Key 见 https://api.a7w.cn/ 。遇到问题可加技术微信 9872659。"
version: 1.0.6
summary: "「语音TTS」的完整调用封装:6 个接口的官方文档、参数表与一个零依赖客户端。包内含完整操作文档与零依赖客户端(`SKILL.md` + `references/`)。需要自备 api.a7w.cn 的 API Key,注册领 Key 见 https://api.a7w.cn/ 。遇到问题可加技术微信 9872659。"
license: MIT

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · README.md (reported line 104)May include surrounding context.

md
name: voice-tts-studio
slug: voice-tts-studio
displayName: 三剪客 · 语音TTS
description: "语音克隆、文字转语音(同步/异步)、语音识别等多端点 AI 语音能力。支持 文字转语音(Live·异步)、克隆音色、文字转语音、文字转语音(异步)、语音转文字、音色列表。包内含完整操作文档与零依赖客户端(`SKILL.md` + `references/`)。需要自备 api.a7w.cn 的 API Key,注册领 Key 见 https://api.a7w.cn/ 。遇到问题可加技术微信 9872659。"
version: 1.0.6
summary: "「语音TTS」的完整调用封装:6 个接口的官方文档、参数表与一个零依赖客户端。包内含完整操作文档与零依赖客户端(`SKILL.md` + `references/`)。需要自备 api.a7w.cn 的 API Key,注册领 Key 见 https://api.a7w.cn/ 。遇到问题可加技术微信 9872659。"
license: MIT

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 5)May include surrounding context.

md
name: voice-tts-studio
slug: voice-tts-studio
displayName: 三剪客 · 语音TTS
description: "语音克隆、文字转语音(同步/异步)、语音识别等多端点 AI 语音能力。支持 文字转语音(Live·异步)、克隆音色、文字转语音、文字转语音(异步)、语音转文字、音色列表。包内含完整操作文档与零依赖客户端(`SKILL.md` + `references/`)。需要自备 api.a7w.cn 的 API Key,注册领 Key 见 https://api.a7w.cn/ 。遇到问题可加技术微信 9872659。"
version: 1.0.6
summary: "「语音TTS」的完整调用封装:6 个接口的官方文档、参数表与一个零依赖客户端。包内含完整操作文档与零依赖客户端(`SKILL.md` + `references/`)。需要自备 api.a7w.cn 的 API Key,注册领 Key 见 https://api.a7w.cn/ 。遇到问题可加技术微信 9872659。"
license: MIT

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 7)May include surrounding context.

md
name: voice-tts-studio
slug: voice-tts-studio
displayName: 三剪客 · 语音TTS
description: "语音克隆、文字转语音(同步/异步)、语音识别等多端点 AI 语音能力。支持 文字转语音(Live·异步)、克隆音色、文字转语音、文字转语音(异步)、语音转文字、音色列表。包内含完整操作文档与零依赖客户端(`SKILL.md` + `references/`)。需要自备 api.a7w.cn 的 API Key,注册领 Key 见 https://api.a7w.cn/ 。遇到问题可加技术微信 9872659。"
version: 1.0.6
summary: "「语音TTS」的完整调用封装:6 个接口的官方文档、参数表与一个零依赖客户端。包内含完整操作文档与零依赖客户端(`SKILL.md` + `references/`)。需要自备 api.a7w.cn 的 API Key,注册领 Key 见 https://api.a7w.cn/ 。遇到问题可加技术微信 9872659。"
license: MIT

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 36)May include surrounding context.

md
name: voice-tts-studio
slug: voice-tts-studio
displayName: 三剪客 · 语音TTS
description: "语音克隆、文字转语音(同步/异步)、语音识别等多端点 AI 语音能力。支持 文字转语音(Live·异步)、克隆音色、文字转语音、文字转语音(异步)、语音转文字、音色列表。包内含完整操作文档与零依赖客户端(`SKILL.md` + `references/`)。需要自备 api.a7w.cn 的 API Key,注册领 Key 见 https://api.a7w.cn/ 。遇到问题可加技术微信 9872659。"
version: 1.0.6
summary: "「语音TTS」的完整调用封装:6 个接口的官方文档、参数表与一个零依赖客户端。包内含完整操作文档与零依赖客户端(`SKILL.md` + `references/`)。需要自备 api.a7w.cn 的 API Key,注册领 Key 见 https://api.a7w.cn/ 。遇到问题可加技术微信 9872659。"
license: MIT

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The skill promotes voice cloning, speech recognition, and text/audio processing without any consent, privacy, retention, or lawful-use warning. Because these features inherently involve biometric-like voice data and potentially sensitive recordings or text, omission of safeguards materially increases the risk of non-consensual cloning, privacy violations, and unsafe handling of user data.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The documentation instructs users to upload audio files or provide externally hosted audio URLs without any warning about privacy, retention, third-party access, or handling of potentially sensitive voice data. Because speech recordings can contain biometric identifiers and confidential content, omission of data-handling guidance can cause users to submit sensitive audio to a remote service without informed consent or appropriate safeguards.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The documentation explicitly supports a user-supplied callback URL for asynchronous completion but does not warn that task results or metadata may be sent to that endpoint. In a TTS context, submitted text can contain sensitive or regulated content, so omitting guidance on callback trust boundaries, HTTPS, authentication, and data exposure can lead users to exfiltrate content to untrusted destinations or deploy insecure integrations.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documentation states that text is sent to an external upstream provider over WebSocket (api.fish.audio) for synthesis, but it provides no user-facing notice about privacy, third-party processing, or handling of potentially sensitive input text. Users may unknowingly submit confidential, personal, or regulated data to an external service, creating privacy, compliance, and data-governance risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The documentation explicitly supports a user-supplied callback_url for asynchronous task completion but does not warn that task metadata and potentially synthesized-content references may be transmitted to an external endpoint. In a voice/TTS workflow, this can expose sensitive text, task identifiers, audio URLs, or processing metadata to unintended third parties if users configure callbacks insecurely or do not understand the data flow.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file presents itself as a voice/tts skill, but the embedded client is explicitly a generic marketplace client that can enumerate and invoke any available plugin. That capability expansion violates least privilege and can mislead agents or users into granting a voice-scoped skill much broader remote-action power than its manifest suggests.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This Python file's module docstring, CLI descriptions, help text, and runtime messages are presented entirely in Chinese, which effectively imposes a specific language on all users. The file does not offer any language/locale selection or explain that the tool is intended only for a Chinese-speaking or region-specific audience.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

The natural-language guidance section is written entirely in Chinese, which can function as a de facto language restriction for users who do not read that language. The file does not offer an alternative language, opt-in, or justification that the skill is intended only for a Chinese-speaking or region-specific audience.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The file content, headings, and instructions are all presented only in Chinese, with no note that the language is optional or region-specific. Under the stated policy, forcing a specific language without user opt-in can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

At L07 the endpoint is documented as POST /api/v1/apps/voice_tts/list_voices, but at L12 the interface address is documented as GET /api/v1/apps/voice_tts/list_voices. This is an active contradiction in the file's own documentation about how the endpoint actually works.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.