Back to skill

Security audit

火山引擎豆包语音播客

Security checks for vulnerabilities and agentic risk

Overview

This skill appears to be a real Volcano Engine podcast generator, but it needs Review because it handles API credentials, sends user text to an external service, and contains unsafe file/endpoint handling.

Review before installing. Use this only with a trusted Volcano Engine account, avoid submitting secrets or regulated data as podcast topics, keep the endpoint fixed to the official Volcano Engine service, and run it under an unprivileged account with write access limited to an intended output directory. Treat the QQ bot helper as higher risk until output filenames are sanitized and config/media path behavior is clearly documented.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/generate_podcast.py:158
Finding

Authentication Credentials Can Be Sent to an Untrusted WebSocket Endpoint

Content
View full analysis

Vulnerability Details

File Location: scripts/generate_podcast.py, lines 158-166, 225-231, and 268-272
Vulnerability Type: Unrestricted credential-bearing endpoint configuration
Risk Level: High

Vulnerable Code

python
def __init__(
    self,
    appid: str,
    access_token: str,
    app_key: str = "aGjiRDfUWi",
    resource_id: str = DEFAULT_RESOURCE_ID,
    endpoint: str = ENDPOINT,
):
    self.appid = appid
    self.access_token = access_token
    self.app_key = app_key
    self.resource_id = resource_id
    self.endpoint = endpoint
python
headers = {
    "X-Api-App-Id": self.appid,
    "X-Api-App-Key": self.app_key,
    "X-Api-Access-Key": self.access_token,
    "X-Api-Resource-Id": self.resource_id,
    "X-Api-Connect-Id": str(uuid.uuid4()),
}
python
websocket = await websockets.connect(
    self.endpoint,
    additional_headers=headers
)

Technical Analysis

PodcastGenerator exposes endpoint as a caller-controlled constructor argument. The value is passed directly to websockets.connect() without validating its scheme, hostname, port, or relationship to the expected Volcengine service.

The connection includes the Volcengine App ID, App Key, access token, resource ID, and connection ID as HTTP headers. After the connection is established, the requested podcast text is also sent in the session payload. Consequently, an attacker-controlled endpoint can collect both authentication material and user-supplied content.

WebSocket TLS does not mitigate this issue when the attacker owns a domain with a valid certificate. TLS would secure the connection to the attacker rather than verify that the recipient is the intended Volcengine service.

Attack Path

  1. An attacker influences application code or configuration that constructs PodcastGenerator.
  2. The attacker supplies an endpoint such as wss://attacker.example/ws.

...[truncated 969 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove endpoint configurability if custom PodcastTTS endpoints are not a required feature.
  2. Otherwise, parse and normalize the endpoint before connecting and enforce:
    • The wss scheme.
    • An explicit hostname allowlist, such as openspeech.bytedance.com.
    • The expected port.
    • No URL user information.
    • No IP literals or loopback, private, link-local, and reserved addresses.
  3. Do not forward authentication headers across redirects or endpoint changes.
  4. Separate credential-bearing production connections from any custom endpoint or test functionality.
  5. Fail closed when validation is inconclusive.
  6. Rotate any credentials that may previously have been used with an untrusted endpoint.
  7. Add tests proving that alternate domains, deceptive subdomains, plain ws URLs, IP addresses, and malformed URLs are rejected.

T09 · Insecure Skill Coding Practices

Error
Location
scripts/kamei_podcast.py:69
Finding

Unsanitized Output Name Enables Path Traversal and File Overwrite

Content
View full analysis

Vulnerability Details

File Location: scripts/kamei_podcast.py, lines 69-70, 91-94, and 120
Vulnerability Type: Path traversal and unsafe file overwrite
Risk Level: High

Vulnerable Code

python
# 创建输出目录
output_dir = Path(f"/tmp/podcast-{output_name}")
output_dir.mkdir(parents=True, exist_ok=True)
python
# 复制到发送目录
src_file = Path(result["final_files"][0])
dst_file = SEND_DIR / f"{output_name}.mp3"

shutil.copy(src_file, dst_file)
python
parser.add_argument("-n", "--name", default="podcast", help="输出文件名")

Technical Analysis

The command-line-controlled output_name is incorporated into filesystem paths without validation. It may contain path separators, .. traversal components, or an absolute path.

The most direct issue occurs when constructing dst_file. In pathlib, joining a base path with an absolute second operand discards the base path. Relative values containing sufficient ../ components can similarly escape SEND_DIR. The subsequent shutil.copy() operation replaces an existing destination file by default.

The temporary output directory also embeds the untrusted name in a path. Although its fixed /tmp/podcast- prefix limits how an absolute value is interpreted in that particular expression, traversal components can still produce unexpected directories outside the intended per-podcast location after path normalization.

Attack Path

  1. An attacker obtains control over the --name argument or the output_name parameter exposed by generate_podcast().
  2. The attacker supplies a traversal value such as ../../../../tmp/attacker-output or an absolute destination name.
  3. Podcast generation completes and produces a source MP3 file.
  4. SEND_DIR / f"{output_name}.mp3" resolves outside the intended QQBot download directory.
  5. shutil.copy() writes the generated audio to the attacker-selected .mp3 path and overwrites an existing file if ...[truncated 894 chars]
Remediation
View remediation

Remediation Suggestions

  1. Treat output_name strictly as a basename rather than a path.
  2. Enforce an allowlist such as ^[A-Za-z0-9_-]{1,64}$.
  3. Explicitly reject:
    • Absolute paths.
    • / and \ separators.
    • . and .. path components.
    • Control characters and platform-specific reserved names.
  4. Resolve the final destination and verify containment before writing:
python
import re

if not re.fullmatch(r"[A-Za-z0-9_-]{1,64}", output_name):
    raise ValueError("Invalid output name")

base_dir = SEND_DIR.resolve()
dst_file = (base_dir / f"{output_name}.mp3").resolve()

if dst_file.parent != base_dir:
    raise ValueError("Output path escapes the media directory")
  1. Create SEND_DIR securely and verify that it is not a symbolic link to an unintended location.
  2. Avoid silently replacing existing files. Use a unique generated filename or an exclusive-create workflow, and return an error on collisions.
  3. Generate temporary output with tempfile.TemporaryDirectory() rather than deriving a /tmp directory from user input.
  4. Run the helper under a dedicated, unprivileged service account with write access only to the required media directory.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (13)

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 67)May include surrounding context.

md
| 参数 | 类型 | 必填 | 默认值 | 说明 |
|------|------|------|--------|------|
| appid | str | 是 | - | 应用 ID |
| access_token | str | 是 | - | Access Token |
| app_key | str | 否 | aGjiRDfUWi | App Key |
| resource_id | str | 否 | volc.service_type.10050 | 资源 ID |
| endpoint | str | 否 | wss://openspeech... | WebSocket 端点 |

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 143)May include surrounding context.

md
| 参数 | 类型 | 必填 | 默认值 | 说明 |
|------|------|------|--------|------|
| appid | str | 是 | - | 应用 ID |
| access_token | str | 是 | - | Access Token |
| app_key | str | 否 | aGjiRDfUWi | App Key |
| resource_id | str | 否 | volc.service_type.10050 | 资源 ID |
| endpoint | str | 否 | wss://openspeech... | WebSocket 端点 |

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/generate_podcast.py (reported line 451)May include surrounding context.

python
| 参数 | 类型 | 必填 | 默认值 | 说明 |
|------|------|------|--------|------|
| appid | str | 是 | - | 应用 ID |
| access_token | str | 是 | - | Access Token |
| app_key | str | 否 | aGjiRDfUWi | App Key |
| resource_id | str | 否 | volc.service_type.10050 | 资源 ID |
| endpoint | str | 否 | wss://openspeech... | WebSocket 端点 |

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill documentation describes capabilities that require environment variable access and local file read/write, but it does not declare any explicit tool scope or permissions. This is dangerous because users and hosting platforms cannot clearly assess or constrain what the skill needs, increasing the risk of over-privileged execution and unintended access to local data or secrets.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill sends user-provided text to a third-party WebSocket API, but the documentation does not prominently warn users that their input leaves the local environment and is processed by an external provider. This creates a privacy and data-handling risk, especially if users submit sensitive content assuming local-only processing.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

This code sends the user's input text to a remote WebSocket API and includes authentication headers, but there is no explicit user-facing disclosure at the point of execution that the topic text will be transmitted to an external service. While the file has logging for connection status, it does not clearly warn users about the privacy implication of sending their content off-machine.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The script takes arbitrary user-provided text and sends it to an external podcast generation service using configured credentials, but this file provides no user-facing notice or consent boundary for that data transfer. If users provide sensitive or regulated content, it could be unintentionally disclosed to a third party, creating privacy and compliance risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

These functions send payload bytes and session identifiers over a network connection via websocket.send, which is a safety-relevant data transmission operation for code files. Although there is internal logging, there is no user-facing confirmation, warning comment, or docstring explaining that user/session data is transmitted to a remote service.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The documented voice_type options are limited to zh_male and zh_female, and the skill description is entirely Chinese, suggesting a Chinese-language/locale constraint. The file does not explicitly explain this as a region-specific limitation or present it as a user opt-in language choice.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The natural-language interface and option values only expose Chinese voice types such as zh_male and zh_female, which implies a language/locale restriction. The file does not clearly document that this is a China/Chinese-specific tool or present this limitation as an explicit user choice with justification.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
76% confidence
Finding

The manifest describes a content-generation function, but the implementation reaches into ~/.openclaw/config.json and environment variables to obtain credentials. While likely used for API access, this is still an additional capability involving host configuration and secret access that is not reflected in the stated scope.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

The manifest description is limited to generating dual-speaker podcast audio from topic text. This script goes beyond generation by copying the output into /root/.openclaw/media/qqbot/downloads and later emitting a <qqvoice> tag for message delivery, which adds a distribution/integration behavior not stated in the manifest.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The script writes a generated MP3 into /root/.openclaw/media/qqbot/downloads, which is a side effect affecting the local system and downstream delivery path. While the module description mentions sending to the user, there is no explicit warning at the point of execution that a file will be written into this directory before the operation occurs.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.