T09 · Insecure Skill Coding Practices
- Location
mimo_api.py:24- Finding
Custom API endpoint can exfiltrate API credentials and private media in the Python client
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill does what it says by sending selected media to Xiaomi MiMo for analysis, but it also has under-scoped credential and endpoint handling that users should review before installing.
Install only if you are comfortable sending selected images, videos, audio, prompts, and a Xiaomi API key to a remote service. Do not set MIMO_API_ENDPOINT unless you fully trust the destination, and avoid using this skill on sensitive or regulated media without explicit approval.
mimo_api.py:24Custom API endpoint can exfiltrate API credentials and private media in the Python client
mimo_api.sh:16Custom API endpoint can exfiltrate API credentials and private media in the shell client
mimo_api.sh:32HOME-derived file path is interpolated into executable Python source
The request destination is taken from the MIMO_API_ENDPOINT environment variable and used directly in requests.post while attaching the resolved API key in headers and sending user-supplied media/question content. If an attacker can influence the environment, they can redirect traffic to an arbitrary host and exfiltrate the API key plus uploaded local file contents and prompts; in this skill, that is especially sensitive because local image/video/audio files are converted to data URIs and transmitted wholesale.
def call_api(content, max_tokens=65536, timeout=300):
"""调用 MiMo API 并返回结果 / Call MiMo API and return result"""
t0 = time.time()
resp = requests.post(API_URL, headers=get_headers(), json={
"model": MODEL,
"messages": [{"role": "user", "content": content}],
"max_completion_tokens": max_tokens,
Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.
start_time=$(date +%s%N)
local resp
resp=$(curl -s --max-time "$timeout" "$API_URL" \
-H "api-key: $MIMO_API_KEY" \
-H "Content-Type: application/json" \
-d @"$tmpfile")
The skill invokes external scripts and describes capabilities that require shell execution, network access, environment access, and potentially file writes, but it does not declare any tool scope or permissions boundaries. This weakens reviewability and enforcement, making it easier for an agent to use broader capabilities than users or operators expect when handling untrusted media inputs and remote URLs.
The skill explicitly instructs the agent to return stdout directly to the user while the workflow sends user-provided images, videos, and audio to an external API. Without a disclosure and consent step, users may unknowingly transmit sensitive visual, audio, or OCR-extracted data off-platform, creating privacy, compliance, and data handling risks.
The tool sends local media content and the user's question to a remote API, but there is no explicit disclosure or confirmation at the point of transmission. In this skill, local files are base64-encoded and uploaded in full, so a user may unintentionally transmit sensitive images, recordings, or videos off-device.
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
set -euo pipefail
API_URL="${MIMO_API_ENDPOINT:-https://api.xiaomimimo.com/v1/chat/completions}"
MODEL="${MIMO_OMNI_MODEL:-clawm-alpha}"
# ============================================================
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
set -euo pipefail
API_URL="${MIMO_API_ENDPOINT:-https://api.xiaomimimo.com/v1/chat/completions}"
MODEL="${MIMO_OMNI_MODEL:-clawm-alpha}"
# ============================================================
The skill converts local image, video, and audio files to data URIs and sends them, along with the user's prompt, to a remote third-party API. If users believe analysis is local or are not clearly warned, sensitive media, OCR content, and embedded secrets could be transmitted off-device without informed consent.
This curl call transmits request bodies containing user prompts, media content, and the API key to an external endpoint. In the context of a multimodal skill this is expected behavior, but it still creates a real confidentiality boundary crossing and can expose sensitive local content to a remote service.
start_time=$(date +%s%N)
local resp
resp=$(curl -s --max-time "$timeout" "$API_URL" \
-H "api-key: $MIMO_API_KEY" \
-H "Content-Type: application/json" \
-d @"$tmpfile")
The natural-language instructions force Windows users to use the Python path and disallow bash/curl, while prescribing preferred tooling for macOS/Linux. This is a platform/tooling policy constraint presented as mandatory behavior without offering user choice or explaining a compliance or safety reason.
The manifest describes a multimodal analysis skill for images, video, and audio, but does not indicate any need to inspect local user configuration files. _resolve_api_key() falls back to reading ~/.openclaw/openclaw.json, which expands the skill’s access to local credential material beyond the obvious requirements of media analysis itself.
The manifest describes a multimodal analysis skill for images, video, and audio, but the code also inspects ~/.openclaw/openclaw.json to extract an API key. Accessing unrelated local configuration and credentials is not part of the user-visible multimodal analysis purpose and is not explicitly declared in the manifest.
No suspicious patterns detected.