Back to skill

Security audit

mimo-omni

Security checks for vulnerabilities and agentic risk

Overview

This skill does what it says by sending selected media to Xiaomi MiMo for analysis, but it also has under-scoped credential and endpoint handling that users should review before installing.

Install only if you are comfortable sending selected images, videos, audio, prompts, and a Xiaomi API key to a remote service. Do not set MIMO_API_ENDPOINT unless you fully trust the destination, and avoid using this skill on sensitive or regulated media without explicit approval.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
mimo_api.py:24
Finding

Custom API endpoint can exfiltrate API credentials and private media in the Python client

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
mimo_api.sh:16
Finding

Custom API endpoint can exfiltrate API credentials and private media in the shell client

Content
View full analysis
"$tmpfile" local start_time start_time=$(date +%s%N) local resp resp=$(curl -s --max-time "$timeout" "$API_URL" \ -H "api-key: $MIMO_API_KEY" \ -H "Content-Type: application/json" \ -d @"$tmpfile") ``` Local files are incorporated into the request as base64 data URIs: ```bash local b64 b64=$(base64 -w 0 "$src") echo "data:${mime};base64,${b64}" ``` ### Technical Analysis The shell client trusts `MIMO_API_ENDPOINT` without validating the URL scheme or destination hostname. It then sends the MiMo API key in an HTTP header and submits a JSON body that can include the complete base64-encoded contents of user-selected local media. Encoding local media is necessary for the documented API operation, but the minimum privilege and trust requirement is to send that content only to the intended service. An unrestricted environment-variable override permits a different server to receive both the credential and private content. Plaintext HTTP is also accepted. The pre-scan warning concerning `curl | bash` is not confirmed. The audited shell script uses `curl` to submit JSON to an API and does not pipe downloaded content into a shell. The actual risk is unrestricted credential-bearing network transmission, not remote script execution. ### Attack Path 1. An attacker controls or influences `MIMO_API_ENDPOINT` in the environment of the Skill process. 2. The endpoint is changed to an attacker-operated server. 3. A user or agent invokes `mimo_api.sh` wi ...[truncated 830 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
mimo_api.sh:32
Finding

HOME-derived file path is interpolated into executable Python source

Content
View full analysis
/dev/null) || true fi fi [[ -n "${MIMO_API_KEY:-}" ]] || die "未找到 MiMo API 密鑰 / No API key found。请设置环境变量 MIMO_API_KEY,或配置 ~/.openclaw/openclaw.json / Set MIMO_API_KEY or configure ~/.openclaw/openclaw.json" } ``` ### Technical Analysis The `_openclaw` path is derived from the `HOME` environment variable and inserted directly into a Python program passed to `python3 -c`. Shell quoting protects the shell command from ordinary shell metacharacters, but it does not make the value safe as Python source. A path containing a single quote and suitable Python syntax can terminate the string literal used by `open('$_openclaw')` and alter the generated program. Exploitation also requires the `[[ -f "$_openclaw" ]]` check to succeed, meaning the attacker must be able to arrange a matching filesystem path in addition to controlling `HOME`. This is an avoidable code-injection boundary. Paths should be supplied as data through `sys.argv` or an environment variable, not embedded into executable source. ### Attack Path 1. An attacker gains control over the `HOME` value used to launch the shell client and can create a corresponding crafted directory and configuration-file path. 2. The attacker selects a path containing characters that break out of the Python string literal and introduce attacker-controlled Python statements. 3. `check_key()` confirms that the crafted path exists. 4. The shell expands ` ...[truncated 716 chars]
Remediation
View remediation
/dev/null ) || true ``` Additional hardening measures: 1. Validate that `HOME` is an absolute, expected user-home path when operating in a privileged or automated context. 2. Avoid dynamically generating source code from environment-derived values. 3. Catch `json.JSONDecodeError` and `OSError` to handle malformed or inaccessible configuration files safely. 4. Ensure the configuration file has restrictive ownership and permissions before reading credentials from it. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (12)

Tainted flow: 'API_URL' from os.environ.get (line 25, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
95% confidence
Finding

The request destination is taken from the MIMO_API_ENDPOINT environment variable and used directly in requests.post while attaching the resolved API key in headers and sending user-supplied media/question content. If an attacker can influence the environment, they can redirect traffic to an arbitrary host and exfiltrate the API key plus uploaded local file contents and prompts; in this skill, that is especially sensitive because local image/video/audio files are converted to data URIs and transmitted wholesale.

Content

Scanner excerpt · mimo_api.py (reported line 81)May include surrounding context.

python
def call_api(content, max_tokens=65536, timeout=300):
    """调用 MiMo API 并返回结果 / Call MiMo API and return result"""
    t0 = time.time()
    resp = requests.post(API_URL, headers=get_headers(), json={
        "model": MODEL,
        "messages": [{"role": "user", "content": content}],
        "max_completion_tokens": max_tokens,

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · mimo_api.sh (reported line 93)May include surrounding context.

sh
start_time=$(date +%s%N)

    local resp
    resp=$(curl -s --max-time "$timeout" "$API_URL" \
        -H "api-key: $MIMO_API_KEY" \
        -H "Content-Type: application/json" \
        -d @"$tmpfile")

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill invokes external scripts and describes capabilities that require shell execution, network access, environment access, and potentially file writes, but it does not declare any tool scope or permissions boundaries. This weakens reviewability and enforcement, making it easier for an agent to use broader capabilities than users or operators expect when handling untrusted media inputs and remote URLs.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill explicitly instructs the agent to return stdout directly to the user while the workflow sends user-provided images, videos, and audio to an external API. Without a disclosure and consent step, users may unknowingly transmit sensitive visual, audio, or OCR-extracted data off-platform, creating privacy, compliance, and data handling risks.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The tool sends local media content and the user's question to a remote API, but there is no explicit disclosure or confirmation at the point of transmission. In this skill, local files are base64-encoded and uploaded in full, so a user may unintentionally transmit sensitive images, recordings, or videos off-device.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · mimo_api.py (reported line 25)May include surrounding context.

python
set -euo pipefail

API_URL="${MIMO_API_ENDPOINT:-https://api.xiaomimimo.com/v1/chat/completions}"
MODEL="${MIMO_OMNI_MODEL:-clawm-alpha}"

# ============================================================

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · mimo_api.sh (reported line 16)May include surrounding context.

sh
set -euo pipefail

API_URL="${MIMO_API_ENDPOINT:-https://api.xiaomimimo.com/v1/chat/completions}"
MODEL="${MIMO_OMNI_MODEL:-clawm-alpha}"

# ============================================================

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill converts local image, video, and audio files to data URIs and sends them, along with the user's prompt, to a remote third-party API. If users believe analysis is local or are not clearly warned, sensitive media, OCR content, and embedded secrets could be transmitted off-device without informed consent.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
88% confidence
Finding

This curl call transmits request bodies containing user prompts, media content, and the API key to an external endpoint. In the context of a multimodal skill this is expected behavior, but it still creates a real confidentiality boundary crossing and can expose sensitive local content to a remote service.

Content

Scanner excerpt · mimo_api.sh (reported line 93)May include surrounding context.

sh
start_time=$(date +%s%N)

    local resp
    resp=$(curl -s --max-time "$timeout" "$API_URL" \
        -H "api-key: $MIMO_API_KEY" \
        -H "Content-Type: application/json" \
        -d @"$tmpfile")

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

The natural-language instructions force Windows users to use the Python path and disallow bash/curl, while prescribing preferred tooling for macOS/Linux. This is a platform/tooling policy constraint presented as mandatory behavior without offering user choice or explaining a compliance or safety reason.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The manifest describes a multimodal analysis skill for images, video, and audio, but does not indicate any need to inspect local user configuration files. _resolve_api_key() falls back to reading ~/.openclaw/openclaw.json, which expands the skill’s access to local credential material beyond the obvious requirements of media analysis itself.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The manifest describes a multimodal analysis skill for images, video, and audio, but the code also inspects ~/.openclaw/openclaw.json to extract an API key. Accessing unrelated local configuration and credentials is not part of the user-visible multimodal analysis purpose and is not explicitly declared in the manifest.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.