Back to skill

Security audit

Minimax Image Understanding

Security checks for vulnerabilities and agentic risk

Overview

This image-understanding skill is mostly purpose-aligned, but it needs Review because it can upload any chosen local file to external APIs and can send MiniMax data and credentials to an unvalidated custom host.

Install only if you are comfortable sending selected images and prompts to MiniMax, OpenAI, or Anthropic. Do not use it on confidential screenshots, credentials, regulated documents, or private files unless your policy permits that upload. Avoid setting MINIMAX_API_HOST to anything other than the intended trusted MiniMax endpoint, and prefer a version that validates image files, limits file size and path scope, and confirms before upload.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/understand_image.py:13
Finding

Configurable MiniMax Endpoint Can Exfiltrate Image Data and API Credentials

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/understand_image.py:18
Finding

Missing Image Validation Allows Arbitrary Local File Transmission

Content
View full analysis
2 else "minimax" prompt = sys.argv[3] if len(sys.argv) > 3 else "直接描述这张图片的业务含义和数据内容,不要罗列元素位置关系" result = understand_image(image_path, model, prompt) ``` ### Technical Analysis The documentation declares support for PNG, JPG, JPEG, GIF, and WebP images, but the implementation accepts any readable filesystem path. It does not verify that the target is a regular file, enforce an approved path scope, decode the content as an image, validate its actual MIME type, or impose a maximum size. The MiniMax branch derives a media type from the filename suffix, but this does not validate the underlying content. The OpenAI and Anthropic branches label all input as PNG regardless of its actual format. In every branch, the complete file is read into memory, Base64-encoded, and sent to an external API. In a direct command-line context, the caller intentionally chooses the path. In an agent context, however, arguments can be derived from untrusted requests or model decisions. Without validation or confirmation, a manipulated invocation can cause files outside the intended image scope to be ...[truncated 1592 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (17)

Tainted flow: 'headers' from os.environ.get (line 93, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/understand_image.py (reported line 75)May include surrounding context.

python
"max_tokens": 1000
    }
    
    resp = requests.post("https://api.openai.com/v1/chat/completions", headers=headers, json=payload, timeout=60)
    data = resp.json()
    
    if "choices" in data and data["choices"]:

Tainted flow: 'headers' from os.environ.get (line 93, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/understand_image.py (reported line 113)May include surrounding context.

python
]
    }
    
    resp = requests.post("https://api.anthropic.com/v1/messages", headers=headers, json=payload, timeout=60)
    data = resp.json()
    
    if "content" in data:

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill documents code execution that can access environment variables, invoke a Python script, and send data over the network, but it does not declare any explicit tool scope or permissions. This weakens reviewability and least-privilege controls, making it easier for an agent or operator to invoke capabilities beyond what is clearly disclosed.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill instructs users to submit local images to external providers such as MiniMax, OpenAI, or Anthropic, but gives no warning that screenshots, documents, or photos may contain sensitive data that will leave the local environment. In this context, the skill is specifically intended for screenshots, charts, and document photos, which often contain confidential business or personal information, so omission of a privacy warning materially increases the risk of inadvertent data exfiltration.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/understand_image.py (reported line 38)May include surrounding context.

python
"-d", json.dumps(payload)
    ]
    
    result = subprocess.run(cmd, capture_output=True, text=True, timeout=35)
    data = json.loads(result.stdout)
    
    if data.get("base_resp", {}).get("status_code") == 0:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The MiniMax path sends the full image to an external API via curl with no disclosure or confirmation, creating a direct path for unintended data sharing. This is especially risky in the skill context because screenshots and document photos often contain secrets, internal metrics, customer data, or other sensitive business information.

Content

No source excerpt is available for this finding.

Tainted flow: 'cmd' from os.environ.get (line 30, credential/environment) → subprocess.run (code execution)

Medium
Category
Data Flow
Confidence
73% confidence
Finding

The command arguments include API_HOST sourced from an environment variable, so the script can be induced to send the full image payload and bearer token to an attacker-controlled endpoint. This is not classic code execution because shell expansion is not used, but it is still a real SSRF/data-exfiltration style issue when untrusted environment configuration is possible.

Content

Scanner excerpt · scripts/understand_image.py (reported line 38)May include surrounding context.

python
"-d", json.dumps(payload)
    ]
    
    result = subprocess.run(cmd, capture_output=True, text=True, timeout=35)
    data = json.loads(result.stdout)
    
    if data.get("base_resp", {}).get("status_code") == 0:

External Transmission

Medium
Category
Data Exfiltration
Confidence
70% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/understand_image.py (reported line 75)May include surrounding context.

python
"max_tokens": 1000
    }
    
    resp = requests.post("https://api.openai.com/v1/chat/completions", headers=headers, json=payload, timeout=60)
    data = resp.json()
    
    if "choices" in data and data["choices"]:

External Transmission

Medium
Category
Data Exfiltration
Confidence
70% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/understand_image.py (reported line 113)May include surrounding context.

python
"max_tokens": 1000
    }
    
    resp = requests.post("https://api.openai.com/v1/chat/completions", headers=headers, json=payload, timeout=60)
    data = resp.json()
    
    if "choices" in data and data["choices"]:

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/understand_image.py (reported line 75)May include surrounding context.

python
"max_tokens": 1000
    }
    
    resp = requests.post("https://api.openai.com/v1/chat/completions", headers=headers, json=payload, timeout=60)
    data = resp.json()
    
    if "choices" in data and data["choices"]:

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/understand_image.py (reported line 75)May include surrounding context.

python
"max_tokens": 1000
    }
    
    resp = requests.post("https://api.openai.com/v1/chat/completions", headers=headers, json=payload, timeout=60)
    data = resp.json()
    
    if "choices" in data and data["choices"]:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The function base64-encodes the supplied image and transmits it to OpenAI without any user-facing disclosure, consent check, or sensitivity gating. For a skill designed to process screenshots, charts, and document photos, this can expose confidential business data, personal data, credentials, or regulated content to a third-party processor unexpectedly.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/understand_image.py (reported line 113)May include surrounding context.

python
]
    }
    
    resp = requests.post("https://api.anthropic.com/v1/messages", headers=headers, json=payload, timeout=60)
    data = resp.json()
    
    if "content" in data:

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/understand_image.py (reported line 113)May include surrounding context.

python
]
    }
    
    resp = requests.post("https://api.anthropic.com/v1/messages", headers=headers, json=payload, timeout=60)
    data = resp.json()
    
    if "content" in data:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The Anthropic path uploads raw image content to an external model provider without warning the user that the file leaves the local environment. Because the skill's purpose is image understanding of potentially sensitive business artifacts, the lack of notice and approval materially increases privacy and data-handling risk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The visible user-facing description and instructions are entirely in Chinese, which can impose a language constraint on users without opt-in. The file does not indicate that the skill is region-specific or offer alternative language support.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

The module description and user-facing strings are written in Chinese, and the default prompt also enforces Chinese output behavior, without indicating that other languages are supported or that Chinese is a required locale. Under the policy, forcing a specific language without opt-in can be a natural-language policy concern.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.