Back to skill

Security audit

mimo-v2.5-tts|Xiaomi Official Version

Security checks for vulnerabilities and agentic risk

Overview

This skill is a coherent MiMo text-to-speech integration with optional Feishu voice-message delivery, but users should understand that text, voice samples, and audio can be sent to third-party services.

Install only if you are comfortable sending synthesis text and any voice-cloning samples to Xiaomi MiMo, and generated audio plus Feishu app credentials/recipient IDs to Feishu when using the send script. Use voice cloning only for voices you have permission to use, avoid sensitive text, prefer a virtual environment with pinned dependencies, and confirm the recipient before sending audio messages.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:46
Finding

Unpinned Third-Party Dependency Installation

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:46; also documented in scripts/mimo_tts.py:7, scripts/mimo_tts_voiceclone.py:7, and scripts/mimo_tts_voicedesign.py:7
Vulnerability Type: Supply-chain risk through an unpinned dependency
Risk Level: Medium

Vulnerable Code Snippet:

bash
pip install openai

Technical Analysis

The installation instruction does not specify a reviewed package version or verify package integrity with cryptographic hashes. Consequently, installation results may change over time and can include an unexpected, compromised, or incompatible release of the openai package or one of its transitive dependencies.

The scripts subsequently import and execute this dependency:

python
from openai import OpenAI

Although the package name corresponds to a legitimate dependency, the absence of version constraints and integrity verification weakens reproducibility and supply-chain security. This behavior is not required for the Skill's functionality: the same dependency can be installed at a reviewed, fixed version with verified hashes.

Attack Path

  1. A user follows the documented pip install openai instruction.
  2. The package installer resolves the latest available package and its transitive dependencies at installation time.
  3. A compromised package-index account, malicious dependency release, index redirection, or future compromised version supplies attacker-controlled content.
  4. The package content is installed and later imported by one of the TTS scripts.
  5. Attacker-controlled code executes with the privileges of the user running the installer or TTS script.
  6. The code may access the process environment, including MIMO_API_KEY, as well as submitted text, local voice samples, generated audio, and other files accessible to that user.

Impact Assessment

Successful exploitation could result in arbitrary code execution within the installing or ...[truncated 621 chars]

Remediation
View remediation

Remediation Suggestions

  1. Add a dependency manifest that pins a reviewed version of openai and all transitive dependencies.

  2. Generate and enforce cryptographic hashes for every resolved distribution. For example:

    bash
    python3 -m pip install --require-hashes -r requirements.txt
    
  3. Generate the lock file from a trusted package index and retain it in version control.

  4. Install dependencies in a dedicated virtual environment rather than the system Python environment.

  5. Explicitly use the official package index and prevent unintended fallback to untrusted indexes.

  6. Periodically review and update pinned versions after vulnerability and compatibility testing.

  7. Update all four installation references so users are directed to the locked, hash-verified installation process rather than pip install openai.

  8. Run the Skill under a minimally privileged account and expose only the credentials and local files required for the requested operation.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (16)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The documented behavior extends beyond simple TTS into message delivery to Feishu, credential use, and handling recipient identifiers, but the declared purpose frames the skill as TTS-centric. This mismatch can mislead users and reviewers about data flows and side effects, causing audio or metadata to be transmitted to third-party services unexpectedly.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The documented behavior extends beyond simple TTS into message delivery to Feishu, credential use, and handling recipient identifiers, but the declared purpose frames the skill as TTS-centric. This mismatch can mislead users and reviewers about data flows and side effects, causing audio or metadata to be transmitted to third-party services unexpectedly.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill advertises voice cloning and Feishu audio sending without explicit warnings or safeguards around consent, identity misuse, and third-party sharing. In this context, cloned voices can enable impersonation, and sending generated or cloned audio to Feishu can expose personal or sensitive content outside the local environment without clear user awareness.

Content

No source excerpt is available for this finding.

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · scripts/feishu_send_audio.sh (reported line 56)May include surrounding context.

sh
# ── Step 2: 获取 tenant_access_token ─────────────────────

TOKEN=$(curl -s -X POST 'https://open.feishu.cn/open-apis/auth/v3/tenant_access_token/internal' \
  -H 'Content-Type: application/json' \
  -d "{\"app_id\":\"$FEISHU_APP_ID\",\"app_secret\":\"$FEISHU_APP_SECRET\"}" \
  | python3 -c "import sys,json; print(json.load(sys.stdin)['tenant_access_token'])")

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · scripts/feishu_send_audio.sh (reported line 66)May include surrounding context.

sh
DURATION_MS=$(ffprobe -v quiet -show_entries format=duration -of csv=p=0 "$OPUS_FILE" \
  | python3 -c "import sys; print(int(float(sys.stdin.read().strip())*1000))")

FILE_KEY=$(curl -s -X POST 'https://open.feishu.cn/open-apis/im/v1/files' \
  -H "Authorization: Bearer $TOKEN" \
  -F "file_type=opus" -F "file_name=voice.opus" -F "duration=$DURATION_MS" \
  -F "file=@$OPUS_FILE" \

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill documents shell scripts, Python execution, environment-variable use, file reads/writes, and external API calls, but it declares no explicit tool scope or permissions boundary. That makes the effective capability surface broader than what a reviewer or runtime policy can easily constrain, increasing the risk of over-privileged execution or misuse through the skill interface.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The activation criteria include broad phrases like 'say it out loud' or 'voice reply,' which can match many normal conversations and trigger the skill in contexts the user did not specifically intend. Because the skill can generate files, invoke scripts, and optionally send audio externally, over-broad activation raises the chance of unintended execution and data disclosure.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
70% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/feishu_send_audio.sh (reported line 56)May include surrounding context.

sh
# ── Step 2: 获取 tenant_access_token ─────────────────────

TOKEN=$(curl -s -X POST 'https://open.feishu.cn/open-apis/auth/v3/tenant_access_token/internal' \
  -H 'Content-Type: application/json' \
  -d "{\"app_id\":\"$FEISHU_APP_ID\",\"app_secret\":\"$FEISHU_APP_SECRET\"}" \
  | python3 -c "import sys,json; print(json.load(sys.stdin)['tenant_access_token'])")

External Transmission

Medium
Category
Data Exfiltration
Confidence
70% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/feishu_send_audio.sh (reported line 79)May include surrounding context.

sh
# ── Step 4: 发送语音消息 ─────────────────────────────────

RESULT=$(curl -s -X POST "https://open.feishu.cn/open-apis/im/v1/messages?receive_id_type=$RECEIVE_ID_TYPE" \
  -H "Authorization: Bearer $TOKEN" \
  -H 'Content-Type: application/json' \
  -d "{\"receive_id\":\"$FEISHU_RECEIVE_ID\",\"msg_type\":\"audio\",\"content\":\"{\\\"file_key\\\":\\\"$FILE_KEY\\\"}\"}")

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This code sends args.text and optional args.context to https://api.xiaomimimo.com/v1 for remote processing, which may transmit user-provided content off-device. While the script name and API usage imply TTS, there is no explicit user-facing warning in comments, help text, or output that the text/context will be sent to an external service.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/mimo_tts.py (reported line 45)May include surrounding context.

python
if not api_key:
        print("❌ MIMO_API_KEY 未设置 / is not set", file=sys.stderr)
        sys.exit(1)
    return OpenAI(api_key=api_key, base_url="https://api.xiaomimimo.com/v1")


def encode_voice_file(file_path: str) -> str:

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/mimo_tts_voiceclone.py (reported line 40)May include surrounding context.

python
if not api_key:
        print("❌ MIMO_API_KEY 未设置 / is not set", file=sys.stderr)
        sys.exit(1)
    return OpenAI(api_key=api_key, base_url="https://api.xiaomimimo.com/v1")


def encode_voice_file(file_path: str) -> str:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script sends both the user-provided text and a base64-encoded voice sample to a third-party API for voice cloning without any explicit consent, warning, or privacy gate. Because voiceprints are sensitive biometric data and the skill is designed for cloning a person's voice, this creates a meaningful privacy and impersonation risk if users supply another person's audio or do not understand that the sample leaves the local environment.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
86% confidence
Finding

The script is explicitly configured to transmit data to an external service endpoint at api.xiaomimimo.com. In the context of a TTS skill, remote processing is expected, but it still constitutes a genuine external data-transfer surface because user text and style/context inputs leave the local environment and may include sensitive information.

Content

Scanner excerpt · scripts/mimo_tts_voicedesign.py (reported line 38)May include surrounding context.

python
if not api_key:
        print("❌ MIMO_API_KEY 未设置 / is not set", file=sys.stderr)
        sys.exit(1)
    return OpenAI(api_key=api_key, base_url="https://api.xiaomimimo.com/v1")


def main() -> None:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The script sends both user-provided synthesis text and voice-design context directly to a third-party remote API, but provides no explicit warning, consent prompt, or privacy notice at the point of use. In a TTS skill, those inputs may contain sensitive personal data, private prompts, or identifying voice/persona descriptions, so silent transmission creates a real data-exposure risk even if it is expected functionality.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The natural-language interface and help text are explicitly limited to Chinese/English, and the file presents the skill as operating in those languages only. There is no indication that users can opt into another language or that this locale restriction is required by a documented regional constraint.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.