Back to skill

Security audit

mimo-asr

Security checks for vulnerabilities and agentic risk

Overview

This speech-to-text skill does what it claims, but it uploads audio to a cloud service while disabling HTTPS certificate checks, which creates a real privacy and integrity risk.

Install only if you are comfortable sending audio to the named cloud Gradio/Hugging Face service. Do not use it for confidential, regulated, or private recordings unless the TLS verification issue is fixed and you understand the service's data handling. Prefer removing verify=False before use and pinning dependencies in a reviewed requirements file.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/mimo_asr_api.py:32
Finding

TLS Certificate Verification Disabled for Audio Transcription Requests

Content
View full analysis

Vulnerability Details

File Location: scripts/mimo_asr_api.py, lines 32-53
Vulnerability Type: Improper TLS certificate validation
Risk Level: High

Vulnerable Code

python
def upload_audio(filepath):
    """Upload an audio file to the Gradio server."""
    with open(filepath, 'rb') as f:
        r = requests.post(
            f"{API_BASE}/gradio_api/upload",
            files={'files': f},
            timeout=120,
            verify=False
        )
    r.raise_for_status()
    data = r.json()
    if isinstance(data, list) and len(data) > 0:
        return data[0]
    return data


def call_transcribe(audio_data_url, language_tag):
    """Call the transcription API and return its event ID."""
    payload = {"data": [audio_data_url, None, language_tag]}
    r = requests.post(
        f"{API_BASE}/gradio_api/call/transcribe",
        json=payload,
        timeout=30,
        verify=False
    )
    r.raise_for_status()
    data = r.json()
    return data.get('event_id')


def poll_result(event_id):
    """Poll the SSE endpoint for the transcription result."""
    url = f"{API_BASE}/gradio_api/call/transcribe/{event_id}"
    with requests.get(
        url,
        stream=True,
        timeout=180,
        verify=False
    ) as r:

The original source uses verify=False on the requests at lines 32, 43, and 53.

Technical Analysis

The verify=False option disables validation of the remote server's TLS certificate. Although the URLs use HTTPS, the client does not verify that it is communicating with the legitimate Hugging Face-hosted service.

An attacker capable of intercepting network traffic can present an arbitrary certificate without causing the client to reject the connection. The attacker can consequently observe or modify the upload request, transcription initiation request, and streamed transcription response.

The vulnerability affects ...[truncated 1409 chars]

Remediation
View remediation

Remediation Suggestions

Remove verify=False from all Requests calls and rely on certificate verification by default:

python
r = requests.post(
    f"{API_BASE}/gradio_api/upload",
    files={'files': f},
    timeout=120
)

r = requests.post(
    f"{API_BASE}/gradio_api/call/transcribe",
    json=payload,
    timeout=30
)

with requests.get(url, stream=True, timeout=180) as r:
    r.raise_for_status()
    # Process the response.

Additional hardening should include:

  1. Keep the operating system and Python CA certificate bundle current.
  2. If a private CA is genuinely required, pass the path of a narrowly scoped trusted CA bundle through verify="/path/to/ca-bundle.pem" rather than disabling verification.
  3. Do not suppress TLS warnings as a substitute for certificate validation.
  4. Call raise_for_status() on the polling response before processing streamed data.
  5. Document that audio is transmitted to a third-party cloud service and advise users not to upload sensitive recordings without authorization.

T08 · Insecure Dependencies

Note
Location
SKILL.md:25
Finding

Unpinned Third-Party Dependency Installation

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, line 25
Vulnerability Type: Unconstrained third-party dependency
Risk Level: Low

Vulnerable Code

bash
# Install the only dependency
pip install requests

Technical Analysis

The installation instruction resolves requests without a version constraint or package hash. The installed artifact therefore depends on the user's configured Python package index and the newest version available at installation time.

This reduces build reproducibility and creates supply-chain exposure. If the configured index, package account, package release, or network/package-management environment is compromised, users may install code that was not reviewed with this project. The risk is lower than dependency confusion involving a private package name because requests is an established public package, but the instruction still lacks integrity pinning.

Attack Path

  1. A user follows the documented pip install requests instruction.
  2. The user's Python environment is configured to use a compromised or attacker-controlled package index, or an upstream release is compromised.
  3. Pip resolves the unconstrained dependency to an attacker-controlled artifact.
  4. The artifact executes installation-time behavior or malicious runtime behavior when imported by mimo_asr_api.py.
  5. The malicious dependency runs with the same operating-system privileges as the user invoking pip or the transcription script.

This path depends on compromise or substitution within the package supply chain; the project does not itself host or retrieve a known malicious package.

Impact Assessment

A substituted dependency could execute arbitrary Python code with the privileges of the installing or invoking user. Depending on those privileges, possible impact includes:

  • Reading or modifying files accessible to the user.
  • Accessing audio files supplied to the script.
  • Stealing ...[truncated 294 chars]
Remediation
View remediation

Remediation Suggestions

Replace the unconstrained installation instruction with a reviewed, version-pinned dependency manifest. For stronger integrity guarantees, include hashes:

text
requests==REVIEWED_VERSION \
    --hash=sha256:REVIEWED_DISTRIBUTION_HASH

Install it using hash enforcement:

bash
python -m pip install --require-hashes -r requirements.txt

The exact version and hash should be selected after reviewing and testing a current supported release rather than being guessed. Additional measures should include:

  1. Obtain dependencies from a trusted package index.
  2. Review dependency updates before changing the pinned version.
  3. Use an isolated virtual environment with minimal permissions.
  4. Maintain a lock file or hash-pinned requirements file in version control.
  5. Run dependency vulnerability scanning as part of release maintenance.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (10)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The code’s core function is still speech-to-text transcription, so the general domain matches. However, the declared description specifically says the skill works by calling a Gradio API and does not require a local model or API key. The supplied code does not call any remote API at all; instead, it imports a local MimoAudio class, loads model assets from local filesystem paths, and performs inference locally. This is a material implementation and resource mismatch, not just an internal detail, because the required environment, dependencies, and execution model differ substantially from the declared purpose.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill advertises and documents behavior that requires network access and may write output files, but it does not declare any explicit tool scope or permissions boundaries. This is dangerous because an agent or reviewer cannot easily tell what capabilities the skill will exercise, increasing the chance of unintended data exfiltration of user audio to a third-party service or unexpected file writes.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The invocation guidance is broad enough to trigger on common requests like voice-to-text or extracting conversation content, without clearly warning that audio will be uploaded to an external public service. In context, this increases the risk of the skill being used on sensitive recordings when a user may reasonably expect local or privacy-preserving processing.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The module docstring explicitly says the script performs '离线语音转文字 / Offline speech-to-text' and the code loads model and tokenizer directories from local paths before invoking a local model object. This does not match the manifest description of using a Gradio API with no local model required, indicating a semantic mismatch between the skill's stated implementation and actual behavior.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The file-level docstring states that the script is based on an offline MiMo-V2.5-ASR model, while the skill's declared documentation says it uses a Gradio API and requires no local model. This is an active contradiction about how the skill operates, not merely an omitted detail.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The script uploads user-provided audio to a third-party remote ASR endpoint, but it does not provide a clear privacy or data-transfer warning at runtime before transmitting potentially sensitive voice content. In an ASR skill context, audio often contains personal, confidential, or regulated data, so silent exfiltration to an external service creates a real privacy and compliance risk even if the transfer is part of intended functionality.

Content

No source excerpt is available for this finding.

Unsafe Defaults

Medium
Category
Tool Misuse
Confidence
99% confidence
Finding

Disabling TLS certificate verification with verify=False allows man-in-the-middle interception or tampering of uploaded audio during transit. Because this function sends raw audio files, an attacker on the network path could capture sensitive voice data or substitute malicious responses while the client treats the HTTPS connection as trusted.

Content

Scanner excerpt · scripts/mimo_asr_api.py (reported line 30)May include surrounding context.

python
def upload_audio(filepath):
    """上传音频文件到 Gradio 服务器"""
    with open(filepath, 'rb') as f:
        r = requests.post(f"{API_BASE}/gradio_api/upload", files={'files': f}, timeout=120, verify=False)
    r.raise_for_status()
    data = r.json()
    if isinstance(data, list) and len(data) > 0:

External Transmission

Medium
Category
Data Exfiltration
Confidence
83% confidence
Finding

This request sends audio-derived data and language metadata to an external service, which is expected for cloud transcription but still constitutes a security-relevant data transfer. In this skill, the danger is amplified because the tool is specifically designed to process potentially sensitive recordings, so external transmission should be treated as a real privacy boundary crossing rather than a harmless implementation detail.

Content

Scanner excerpt · scripts/mimo_asr_api.py (reported line 41)May include surrounding context.

python
def call_transcribe(audio_data_url, language_tag):
    """调用 transcribe API,返回 event_id"""
    payload = {"data": [audio_data_url, None, language_tag]}
    r = requests.post(f"{API_BASE}/gradio_api/call/transcribe", json=payload, timeout=30, verify=False)
    r.raise_for_status()
    data = r.json()
    return data.get('event_id')

Unsafe Defaults

Medium
Category
Tool Misuse
Confidence
99% confidence
Finding

The transcription API call also disables TLS verification, so the event creation step can be intercepted or modified by an active attacker. This could expose metadata, redirect the workflow, or inject a forged event_id that causes the client to consume attacker-controlled results.

Content

Scanner excerpt · scripts/mimo_asr_api.py (reported line 41)May include surrounding context.

python
def call_transcribe(audio_data_url, language_tag):
    """调用 transcribe API,返回 event_id"""
    payload = {"data": [audio_data_url, None, language_tag]}
    r = requests.post(f"{API_BASE}/gradio_api/call/transcribe", json=payload, timeout=30, verify=False)
    r.raise_for_status()
    data = r.json()
    return data.get('event_id')

Unsafe Defaults

Medium
Category
Tool Misuse
Confidence
99% confidence
Finding

The SSE polling request disables certificate verification while receiving transcription results, enabling a network attacker to spoof or alter returned transcript data. In a speech-to-text workflow, this threatens both confidentiality and integrity: sensitive results can be exposed, and downstream users may trust manipulated transcript output.

Content

Scanner excerpt · scripts/mimo_asr_api.py (reported line 51)May include surrounding context.

python
"""轮询 SSE 端点获取转录结果"""
    url = f"{API_BASE}/gradio_api/call/transcribe/{event_id}"
    # 用 stream 模式读 SSE
    with requests.get(url, stream=True, timeout=180, verify=False) as r:
        for line in r.iter_lines(decode_unicode=True):
            if not line:
                continue

Static analysis

Detected: suspicious.insecure_tls_verification

HTTPS certificate verification is disabled.

Warn
Code
suspicious.insecure_tls_verification
Location
scripts/mimo_asr_api.py:30