Back to skill

Security audit

豆包语音合成 2.0

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent cloud text-to-speech tool, but it needs Review because credentials and user text can be sent to a configurable endpoint and Windows playback uses an unsafe shell call.

Install only if you are comfortable sending synthesis text to Volcengine/ByteDance and storing a Volcengine token for this skill. Do not set VOLCANO_API_URL unless you fully trust the destination, avoid putting secrets or regulated text into TTS requests, and be cautious using playback on Windows until the shell=True playback path is fixed.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
tts_client.py:67
Finding

Authentication Credentials and User Text Can Be Sent to an Untrusted Configurable Endpoint

Content
View full analysis
Dict[str, str]: """Build request headers.""" headers = { "Content-Type": "application/json", "X-Api-App-Id": self.app_id, "X-Api-Access-Key": self.access_token, "X-Api-Resource-Id": self.resource_id, # Add Authorization header because some API nodes require it "Authorization": f"Bearer; {self.access_token}", } if request_id: headers["X-Api-Request-Id"] = request_id return headers ``` ```python response = SeedTTS2._session.post( self.api_url, headers=headers, json=payload, stream=True, timeout=self.timeout ) ``` ### Technical Analysis The API destination can be supplied through the `api_url` constructor argument or the `VOLCANO_API_URL` environment variable. The code does not validate the URL scheme, hostname, port, or path before attaching sensitive authentication headers and transmitting the synthesis payload. Consequently, the client will send the following information to any configured destination: - Volcengine APP ID - Volcengine Access Token in `X-Api-Access-Key` - The same Access Token in the `Authorization` header - Resource identifier - User-supplied text submitted for speech synthesis - Speaker and audio configuration The endpoint override is also not documented in `SKILL ...[truncated 2002 chars]
Remediation
View remediation
str: parsed = urlparse(url) if ( parsed.scheme != "https" or parsed.hostname ! ...[truncated 596 chars]
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Output HandlingUnvalidated Output Injection, Cross-Context Output, Unbounded Output
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
Findings (30)

Tp4

High
Category
MCP Tool Poisoning
Confidence
90% confidence
Finding

While feature mismatch alone is not always a vulnerability, this finding also notes undeclared local file writing, audio playback, credential reading, and external network requests. Hidden or under-declared side effects reduce informed consent and can cause users to expose secrets or transmit data externally without realizing the operational scope.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
90% confidence
Finding

While feature mismatch alone is not always a vulnerability, this finding also notes undeclared local file writing, audio playback, credential reading, and external network requests. Hidden or under-declared side effects reduce informed consent and can cause users to expose secrets or transmit data externally without realizing the operational scope.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 47)May include surrounding context.

md
1. 登录 [火山引擎控制台](https://console.volcengine.com/speech)
2. 进入「语音合成」→「应用管理」
3. 创建应用或选择已有应用
4. 获取 **APP ID** 和 **Access Token**

### 2. 配置凭证

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 55)May include surrounding context.

md
1. 登录 [火山引擎控制台](https://console.volcengine.com/speech)
2. 进入「语音合成」→「应用管理」
3. 创建应用或选择已有应用
4. 获取 **APP ID** 和 **Access Token**

### 2. 配置凭证

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 226)May include surrounding context.

md
1. 登录 [火山引擎控制台](https://console.volcengine.com/speech)
2. 进入「语音合成」→「应用管理」
3. 创建应用或选择已有应用
4. 获取 **APP ID** 和 **Access Token**

### 2. 配置凭证

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · tts_client.py (reported line 61)May include surrounding context.

python
1. 登录 [火山引擎控制台](https://console.volcengine.com/speech)
2. 进入「语音合成」→「应用管理」
3. 创建应用或选择已有应用
4. 获取 **APP ID** 和 **Access Token**

### 2. 配置凭证

Credential Access

High
Category
Privilege Escalation
Confidence
72% confidence
Finding

The documentation recommends storing the Access Token in a local JSON config file under the user's home directory without discussing file-permission hardening or secret-management risks. Plaintext credential storage in config files can expose tokens to other local users, backups, sync services, or accidental commits.

Content

Scanner excerpt · SKILL.md (reported line 70)May include surrounding context.

md
"seedtts2": {
        "env": {
          "VOLCANO_APP_ID": "你的 APP ID",
          "VOLCANO_ACCESS_TOKEN": "你的 Access Token"
        }
      }
    }

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · docs/troubleshooting.md (reported line 22)May include surrounding context.

方式一:环境变量(推荐)

bash
export VOLCANO_APP_ID="你的 APP ID"
export VOLCANO_ACCESS_TOKEN="你的 Access Token"

方式二:OpenClaw 配置

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · docs/troubleshooting.md (reported line 33)May include surrounding context.

方式一:环境变量(推荐)

bash
export VOLCANO_APP_ID="你的 APP ID"
export VOLCANO_ACCESS_TOKEN="你的 Access Token"

方式二:OpenClaw 配置

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · docs/troubleshooting.md (reported line 61)May include surrounding context.

方式一:环境变量(推荐)

bash
export VOLCANO_APP_ID="你的 APP ID"
export VOLCANO_ACCESS_TOKEN="你的 Access Token"

方式二:OpenClaw 配置

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · docs/troubleshooting.md (reported line 70)May include surrounding context.

方式一:环境变量(推荐)

bash
export VOLCANO_APP_ID="你的 APP ID"
export VOLCANO_ACCESS_TOKEN="你的 Access Token"

方式二:OpenClaw 配置

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · docs/troubleshooting.md (reported line 77)May include surrounding context.

方式一:环境变量(推荐)

bash
export VOLCANO_APP_ID="你的 APP ID"
export VOLCANO_ACCESS_TOKEN="你的 Access Token"

方式二:OpenClaw 配置

Unvalidated Output Injection

High
Category
Output Handling
Confidence
95% confidence
Finding

Model output is used without validation or sanitization. Unvalidated output injected into downstream contexts (SQL, shell, HTML) enables injection attacks and arbitrary code execution.

Content

Scanner excerpt · tts_client.py (reported line 326)May include surrounding context.

python
try:
            # 根据系统选择播放器
            if sys.platform == 'darwin':  # macOS
                subprocess.run(['afplay', output], check=True)
            elif sys.platform == 'linux':
                subprocess.run(['aplay', output], check=True)
            elif sys.platform == 'win32':

Unvalidated Output Injection

High
Category
Output Handling
Confidence
95% confidence
Finding

Model output is used without validation or sanitization. Unvalidated output injected into downstream contexts (SQL, shell, HTML) enables injection attacks and arbitrary code execution.

Content

Scanner excerpt · tts_client.py (reported line 328)May include surrounding context.

python
if sys.platform == 'darwin':  # macOS
                subprocess.run(['afplay', output], check=True)
            elif sys.platform == 'linux':
                subprocess.run(['aplay', output], check=True)
            elif sys.platform == 'win32':
                subprocess.run(['start', output], shell=True, check=True)
            else:

Unvalidated Output Injection

High
Category
Output Handling
Confidence
99% confidence
Finding

This is a true unvalidated output injection issue because an attacker-influenced output value reaches a shell-enabled subprocess on Windows. In the skill context, users can supply output paths to say()/synthesize(), making the dangerous sink realistically reachable and more severe than in a closed internal-only utility.

Content

Scanner excerpt · tts_client.py (reported line 330)May include surrounding context.

python
elif sys.platform == 'linux':
                subprocess.run(['aplay', output], check=True)
            elif sys.platform == 'win32':
                subprocess.run(['start', output], shell=True, check=True)
            else:
                print(f"⚠️  未知平台,无法自动播放:{output}", file=sys.stderr)
                return False

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
99% confidence
Finding

This is a true tool-parameter abuse issue because an externally influenced filename is passed into a shell-backed command on Windows. The combination of start and shell=True can let crafted input alter the command being executed, turning a playback helper into an arbitrary command execution vector.

Content

Scanner excerpt · tts_client.py (reported line 330)May include surrounding context.

python
elif sys.platform == 'linux':
                subprocess.run(['aplay', output], check=True)
            elif sys.platform == 'win32':
                subprocess.run(['start', output], shell=True, check=True)
            else:
                print(f"⚠️  未知平台,无法自动播放:{output}", file=sys.stderr)
                return False

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill documentation describes capabilities that inherently require environment-variable access, local file access, network access, and shell/audio playback, but it does not declare any explicit tool scope or permissions. This weakens user visibility and platform enforcement, making it easier for a skill to access sensitive resources without clear consent boundaries.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill instructs users to configure credentials for a cloud TTS service but does not clearly warn that input text will be sent to an external Volcengine service for synthesis. Users may unknowingly submit sensitive or regulated content, creating privacy, compliance, and confidentiality risks.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · docs/troubleshooting.md (reported line 123)May include surrounding context.

Linux:

bash
# 安装 alsa-utils
sudo apt install alsa-utils

# 播放
aplay output.mp3

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · docs/troubleshooting.md (reported line 250)May include surrounding context.

md
"X-Api-Resource-Id": "seed-tts-2.0",
}

response = requests.post(
    "https://openspeech.bytedance.com/api/v3/tts/unidirectional",
    headers=headers,
    json={

External Transmission

Medium
Category
Data Exfiltration
Confidence
70% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · docs/troubleshooting.md (reported line 250)May include surrounding context.

md
"X-Api-Resource-Id": "seed-tts-2.0",
}

response = requests.post(
    "https://openspeech.bytedance.com/api/v3/tts/unidirectional",
    headers=headers,
    json={

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This Python file performs an HTTP POST to a third-party API using the input text as request content and includes authentication headers built from configured credentials. While the module describes TTS functionality, it does not visibly disclose in user-facing output or comments near execution that text and related metadata are transmitted off-device to an external service.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest describes a voice synthesis skill with emotion control and multiple voices, but this method adds OS-level playback by invoking platform-specific executables via subprocess. Generating speech audio is aligned with the stated purpose; spawning local programs to play it is an additional host-execution capability not justified by the manifest description.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · tts_client.py (reported line 326)May include surrounding context.

python
try:
            # 根据系统选择播放器
            if sys.platform == 'darwin':  # macOS
                subprocess.run(['afplay', output], check=True)
            elif sys.platform == 'linux':
                subprocess.run(['aplay', output], check=True)
            elif sys.platform == 'win32':

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · tts_client.py (reported line 328)May include surrounding context.

python
if sys.platform == 'darwin':  # macOS
                subprocess.run(['afplay', output], check=True)
            elif sys.platform == 'linux':
                subprocess.run(['aplay', output], check=True)
            elif sys.platform == 'win32':
                subprocess.run(['start', output], shell=True, check=True)
            else:

Static analysis

No suspicious patterns detected.