Back to skill

Security audit

多语种音频翻译助手

Security checks for vulnerabilities and agentic risk

Overview

This audio translation skill mostly matches its purpose, but it has unsafe execution paths and broad install/network behavior that should be reviewed before use.

Review carefully before installing. Do not use this skill on confidential, personal, regulated, or secret-containing audio unless you are comfortable sending transcript and translated text to external services. Run it only in a contained environment, avoid untrusted audio URLs, install dependencies through a reviewed virtual environment with pinned versions, and fix the Python interpolation bugs before processing attacker-supplied audio or remote translation responses.

Vulnerability Patterns
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (5)

T03 · Remote Payload Retrieval and Execution

Error
Location
scripts/install_deps.sh:13
Finding

Mutable Remote Homebrew Installer Is Recommended for Direct Shell Execution

Content
View full analysis
/dev/null; then echo "Homebrew 未安装,请手动安装:" echo " /bin/bash -c \"\$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)\"" echo "" echo "或者访问: https://brew.sh" exit 1 fi ``` ### Technical Analysis The script prints an instruction that downloads the current `HEAD/install.sh` file from GitHub and executes it immediately through `/bin/bash`. Although `install_deps.sh` does not execute the command automatically, the message explicitly instructs the user to do so. The URL points to Homebrew's official GitHub repository rather than a personal paste service, which reduces source-impersonation risk. However, the `HEAD` resource is mutable, no version or commit is pinned, and no checksum or signature is verified. Consequently, the effective code executed by the user can change after this Skill has been reviewed. Directly passing downloaded content to a shell prevents meaningful inspection and creates a supply-chain execution path. Installing Homebrew may also involve privileged system changes or requests for elevated authorization, which is broader than merely translating audio. ### Attack Path 1. A user runs `scripts/install_deps.sh` on macOS without Homebrew installed. 2. The script prints the `curl`-to-shell installation command. 3. The user copies and executes the suggested command. 4. The mutable installer is downloaded from the current upstream `HEAD`. 5. If the upstream repository, release process, hosting account, or network trust path has been compromised, altered commands execute immediately. 6. The payload obtains the privileges of the invoking user and may obtain additional privileges if the installation process requests administrator authorization. ### Impact Assessment A malic ...[truncated 511 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/translate.sh:247
Finding

Audio Transcription Is Interpolated into Executable Python Source

Content
View full analysis
/dev/null || echo "") ``` ### Technical Analysis `SOURCE_TEXT` originates from Whisper transcription of a user-supplied local audio file or remotely downloaded audio. The value is inserted directly into a Python program passed to `python3 -c`. Triple-quoted Python strings do not safely encode arbitrary content. If transcription produces a sequence containing `'''` followed by valid Python syntax, the string literal can be terminated and additional Python expressions or statements can be introduced. Shell quoting does not protect the Python parser from this injection because expansion occurs before Python receives the generated source. Audio is therefore treated as executable source indirectly. Exploitation depends on Whisper producing sufficiently precise attacker-chosen text, but untrusted transcription must not be embedded into code regardless of the practical reliability of speech-to-text payload construction. ### Attack Path 1. An attacker creates an audio file designed to transcribe into text containing triple quotes and Python syntax. 2. The attacker provides the file directly or hosts it at a URL supplied to the Skill. 3. Whisper processes the audio and writes the resulting text to `source_text.txt`. 4. The shell loads that text into `SOURCE_TEXT`. 5. The translation helper interpolates the transcription into the `python3 -c` source string. 6. If the transcription terminates the triple-quoted literal successfully, injected Python executes under the account running the Skill. ### Impact Assessment Successful exploitation provides arbitrary Python code execution with the privileges of the invoking user. The injected code could access us ...[truncated 264 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/translate.sh:254
Finding

Remote Translation Response and Transcription Are Embedded in Python Code

Content
View full analysis
/dev/null || echo "$SOURCE_TEXT") ``` ### Technical Analysis Both `API_RESULT` and `SOURCE_TEXT` are inserted directly into Python source code. `API_RESULT` is obtained from the external MyMemory service, while `SOURCE_TEXT` is derived from user-controlled audio. JSON is data, not a valid source-code encoding mechanism. A JSON response can legally include quote and escape sequences that alter the generated Python program. A response containing a suitable `'''` sequence can terminate the Python literal and introduce executable statements. The fallback branch independently embeds the untrusted transcription in another triple-quoted literal. TLS protects transport confidentiality and server authentication under normal assumptions, but it does not make externally returned content safe to embed in source. A compromised service, malicious intermediary with trusted certificate capabilities, unexpected service response, or attacker-influenced translation result could reach this sink. ### Attack Path 1. The Skill sends transcribed text and language parameters to the MyMemory endpoint. 2. An attacker influences the returned response through service compromise, response manipulation, or content that causes an unsafe response representation. 3. Alternatively, crafted audio causes malicious syntax to appear in `SOURCE_TEXT`. 4. The shell expands the response and source text into the body of `python3 -c`. 5. A crafted triple-quote sequence terminates the intended string literal. 6. Injected Python executes with the privileges of the user running the Skill. ### Impact Assess ...[truncated 415 chars]
Remediation
View remediation
"$TEMP_DIR/api_result.json" python3 - "$TEMP_DIR/api_result.json" \ "$SOURCE_TEXT_FILE" \ "$TRANSLATED_FILE" <<'PY' import json import pathlib import sys api_file, source_file, output_file = map(pathlib.Path, sys.argv[1:4]) source_text = source_file.read_text(encoding="utf-8") try: with api_file.open(encoding="utf-8") as handle: data = json.load(handle) translated = data["responseData"]["translatedText"] if not isinstance(translated, str): raise TypeError("translatedText is not a string") except (OSError, ValueError, KeyError, TypeError): translated = source_text output_file.write_text(translated, encoding="utf-8") PY ``` ]]>

T08 · Insecure Dependencies

Warning
Location
scripts/requirements.txt:1
Finding

Unpinned Python Dependencies Are Installed Automatically

Content
View full analysis
/dev/null || \ $PYTHON_BIN -m pip install -r requirements.txt ``` They may also be installed during normal translation execution by `scripts/translate.sh:206-208`: ```bash if ! $PYTHON_BIN -c "import faster_whisper" 2>/dev/null; then echo "安装 Python 依赖..." $PYTHON_BIN -m pip install --user -r "$(dirname "$0")/requirements.txt" 2>/dev/null || \ $PYTHON_BIN -m pip install -r "$(dirname "$0")/requirements.txt" fi ``` ### Technical Analysis Neither dependency has a version constraint or cryptographic hash. Each installation can therefore retrieve whatever release and transitive dependency set is current at that time. This makes the installed code non-reproducible and permits the effective behavior of the Skill to change after audit. Python package installation may execute package build logic. Package code also executes when imported during transcription and TTS. Automatic dependency installation during a normal translation run further blurs the boundary between setup and routine operation. The fallback from `--user` installation to a general environment installation can modify the active Python environment. Although the script does not invoke `sudo` for pip, it may change shared environments when run by an account that has permission to do so. ### Attack Path 1. An upstream package, release account, build pipeline, or transitive dependency is compromised. 2. A malicious or unexpectedly vulnerable release becomes the version selected by pip. 3. The user runs the dependency installer or invokes translation without the required module installed. 4. Pip downloads an ...[truncated 551 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/translate.sh:100
Finding

URL Audio Download Permits Requests to Internal and Reserved Network Destinations

Content
View full analysis
/dev/null; then echo -e "${RED}错误: 需要 curl${NC}" exit 1 fi # 获取文件名并安全处理 FILENAME=$(basename "$INPUT_PATH" | sed 's/[^a-zA-Z0-9._-]/_/g') LOCAL_INPUT="$TEMP_DIR/input_${FILENAME}" # 下载文件(带超时和安全选项) echo "正在下载文件..." if ! curl -L --max-time 120 -o "$LOCAL_INPUT" -- "$INPUT_PATH" 2>/dev/null; then echo -e "${RED}错误: 下载失败${NC}" exit 1 fi # 验证下载的文件大小 FILE_SIZE=$(stat -f%z "$LOCAL_INPUT" 2>/dev/null || stat -c%s "$LOCAL_INPUT" 2>/dev/null) if [[ "$FILE_SIZE" -lt 100 ]]; then echo -e "${RED}错误: 下载的文件过小,可能无效${NC}" exit 1 fi ``` ### Technical Analysis The URL check only confirms that the value begins with an HTTP or HTTPS URL whose host starts with an alphanumeric character. It is not a destination whitelist and does not reject loopback, private, link-local, reserved, or cloud metadata addresses. `curl -L` follows redirects without validating each redirect target. An initially public URL can therefore redirect to an internal destination. DNS rebinding or a hostname resolving to a private address can produce the same result. URL-based audio input is part of the declared functionality, but access to arbitrary network destinations exceeds the minimum network privilege needed to retrieve ordinary public audio files. Extension and minimum-size checks occur only after the request and do not prevent access to internal services. ## ...[truncated 1137 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
Findings (17)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The documented behavior understates important security-relevant actions: downloading remote audio, sending recognized text to third-party services, using remote TTS, and potentially modifying the environment via dependency installation. When a skill omits these behaviors, users cannot make informed consent decisions, and sensitive audio-derived content may be exfiltrated unexpectedly.

Content

No source excerpt is available for this finding.

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · scripts/install_deps.sh (reported line 3)May include surrounding context.

sh
#!/bin/bash
# 安装 audio-translator 所需依赖 (安全版本)
# 不再使用 curl | bash

set -e

Chaining Abuse

High
Category
Tool Misuse
Confidence
70% confidence
Finding

Tool calls are chained to bypass individual safety checks or escalate capabilities beyond what any single tool call would allow.

Content

Scanner excerpt · scripts/install_deps.sh (reported line 3)May include surrounding context.

sh
#!/bin/bash
# 安装 audio-translator 所需依赖 (安全版本)
# 不再使用 curl | bash

set -e

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill declares powerful tools (bash, python, curl) and clearly performs file access, shell execution, network download, and file output, but it does not define any explicit tool scope or permission boundaries. This increases the chance of over-broad execution and makes it harder for users or the runtime to constrain what the skill may access or transmit.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill lacks a clear user-facing warning that audio content or derived transcript text may be sent to external translation and speech-synthesis services. In an audio translation context, inputs may contain personal, confidential, or regulated content, so missing disclosure materially raises privacy and compliance risk.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
84% confidence
Finding

The skill sends recognized text to an external translation API over the network, which is a genuine data-exposure vector. Although this is functionally related to translation, the danger depends on the sensitivity of the spoken content; transcripts can include secrets, personal data, or business information, and the current skill text does not describe safeguards, minimization, or consent.

Content

Scanner excerpt · SKILL.md (reported line 79)May include surrounding context.

步骤3: 翻译(MyMemory API)

bash
curl -s "https://api.mymemory.translated.net/get?q=<文本>&langpair=<源>|<目标>"

步骤4: 目标语言语音合成(edge-tts)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script's comments and all user-facing echo messages are written in Chinese, which imposes a specific language on users. Under the policy, locale or language constraints should be optional, user-selectable, or clearly justified as region-specific, and no such opt-in or justification is present here.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · scripts/install_deps.sh (reported line 35)May include surrounding context.

sh
if ! command -v ffmpeg &> /dev/null; then
        echo "安装 FFmpeg..."
        sudo apt update
        sudo apt install -y ffmpeg
    else
        echo "FFmpeg 已安装"

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · scripts/install_deps.sh (reported line 36)May include surrounding context.

sh
if ! command -v ffmpeg &> /dev/null; then
        echo "安装 FFmpeg..."
        sudo apt update
        sudo apt install -y ffmpeg
    else
        echo "FFmpeg 已安装"

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · scripts/translate.sh (reported line 91)May include surrounding context.

sh
# 创建临时目录(安全模式:使用 mktemp -d)
TEMP_DIR=$(mktemp -d)
chmod 700 "$TEMP_DIR"
trap 'rm -rf "$TEMP_DIR"' EXIT

# ========== 检测输入类型并下载/复制 ==========

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script automatically installs Python packages at runtime from requirements.txt using pip. This expands the trust boundary to package indexes and dependency contents during execution, enabling supply-chain compromise or unexpected code execution if dependencies are malicious or tampered with.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The workflow sends recognized speech text to an external translation API and sends translated text to a remote TTS service, but the user is not explicitly warned that their content leaves the local machine. For audio translation, this can expose sensitive spoken content, personal data, or confidential material to third parties without informed consent.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
94% confidence
Finding

The script transmits user-derived text to https://api.mymemory.translated.net for translation. External transmission of potentially sensitive content is security-relevant here because the skill processes speech that may contain private or confidential information, and there is no explicit consent or data-handling notice.

Content

Scanner excerpt · scripts/translate.sh (reported line 252)May include surrounding context.

sh
# 使用 curl 调用翻译 API(避免 Python urllib SSL 问题)
ENCODED_TEXT=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''$SOURCE_TEXT'''))" 2>/dev/null || echo "")
API_URL="https://api.mymemory.translated.net/get?q=${ENCODED_TEXT}&langpair=${SOURCE_LANG}|${TARGET_LANG}"

# 调用 API
API_RESULT=$(curl -s --max-time 30 "$API_URL" 2>/dev/null)

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
97% confidence
Finding

The dependency faster-whisper is unpinned, so future installs may pull a newer release with breaking changes or a compromised/supply-chain-poisoned version. Because this skill processes untrusted audio/URL inputs and likely relies on external ML/audio libraries, uncontrolled dependency drift increases the attack surface and harms reproducibility.

Content

Scanner excerpt · scripts/requirements.txt (reported line 1)May include surrounding context.

text
faster-whisper
edge-tts

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
97% confidence
Finding

The dependency edge-tts is unpinned, allowing installation of whatever version is current at build time. If a malicious or vulnerable upstream release is published, this skill could silently consume it, which is especially relevant for a network-facing translation/TTS workflow that handles external content and may run in automated environments.

Content

Scanner excerpt · scripts/requirements.txt (reported line 2)May include surrounding context.

text
faster-whisper
edge-tts

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

User-facing comments, usage text, status messages, and errors are all presented in Chinese, which effectively forces a specific language experience. There is no indication that this is a region-specific tool or that users can opt into another language.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest describes a speech-translation skill that generates a target-language audio file from URL or local audio input. In addition to that stated output, the code also saves the translated text as a separate .txt file, which is a behavior not mentioned in the manifest description.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.