T05 · Unauthorized Access and Privilege Escalation
- Location
script/invoke_model.py:23- Finding
Undocumented Local File Upload Exceeds the Skill's Declared Scope
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill is a text-to-image wrapper, but it gives the agent broader, under-disclosed authority to run local Python, upload local media files, and send an API token to a configurable endpoint.
Install only if you trust the publisher and the configured API endpoint. Treat prompts and any supplied image or video paths as uploaded data, avoid passing sensitive local files, pin or verify TEAM_BASE_URL before use, and prefer a version that uses the normal OpenClaw runner with explicit upload consent and endpoint validation.
script/invoke_model.py:23Undocumented Local File Upload Exceeds the Skill's Declared Scope
script/invoke_model.py:37Environment-Controlled API Endpoint Can Receive the Bearer Credential and Submitted Content
The request target is derived from the TEAM_BASE_URL environment variable and is used directly for an authenticated POST. If an attacker can influence the environment or skill configuration, they can redirect prompts and any base64-encoded local image/video contents to an attacker-controlled server along with the bearer token, causing credential and data exfiltration.
print(f"Invoking model: {args.model} ...")
try:
response = requests.post(endpoint, headers=headers, json=payload)
response.raise_for_status()
result = response.json()
The polling URL is also built from the same environment-controlled base URL and queried with the Authorization header. This enables continued token disclosure and interaction with an attacker-controlled endpoint after initial submission, expanding the exfiltration window and allowing untrusted response data to drive tool behavior.
while True:
time.sleep(3) # 每3秒查询一次
poll_resp = requests.get(poll_endpoint, headers=headers)
poll_resp.raise_for_status()
poll_data = poll_resp.json()
The skill is ներկայացված as a simple text-to-image capability, but the analyzed behavior indicates broader multimodal processing, local file reading, remote upload, and task polling. This mismatch is dangerous because users and orchestrators may grant trust or invoke the skill under false assumptions, leading to unintended data exposure or execution of more powerful behavior than advertised.
YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).
import os
import sys
import argparse
import base64
import requests
import json
import mimetypes
import time
def parse_args():
parser = argparse.ArgumentParser(description='Invoke Team AI Models via unified API')
parser.add_argument('--model', type=str, required=True, help='The model ID to use (e.g., qwen3_5-plus, sora-2, etc.)')
parser.add_argument('--prompt', type=str, required=True, help='The prompt text to send to the model')
parser.add_argument('--image', type=str, help='Path to the local image file or remote image URL (First frame)')
parser.add_argument('--image-tail', type=str, help='Path to the local image file or remote image URL (Tail/Last frame)')
parser.add_argument('--video', type=str, help='Path to the local video fi
The skill declares executable and external-access capabilities via Python and an API key, but does not define an explicit tool scope such as permissions or allowed-tools. This weakens containment and makes it easier for an agent to invoke code and networked behavior without clear policy boundaries or user-visible constraints.
The invocation description is broad enough that an agent may trigger the skill whenever image generation is loosely relevant, without clear boundaries or exclusions. Overbroad activation increases the chance of unnecessary code execution, API-key use, and transmission of user data to an external service.
The skill does not warn users that invoking it runs a local Python command and uses an API key to contact an external service. Without that disclosure, users cannot meaningfully consent to the execution and data-transfer side effects, which can result in accidental leakage of prompts or attached content.
The repeated instruction that the agent must run the skill whenever the user asks, combined with a hard requirement to execute a direct Python command, pushes the agent toward automatic execution without safety gating. This is especially risky because the skill content includes adversarial-style imperative language that attempts to override safer orchestration paths.
The skill is described as text-to-image, but the code accepts local or remote image/video inputs and can transmit them to a remote API. This scope mismatch is dangerous because users or calling agents may grant trust based on the advertised capability while the implementation handles more sensitive data types than expected.
The script contains explicit logic for image-to-video models, exceeding the stated text-to-image purpose. In an agent-skill context, this increases risk because hidden or broader-than-advertised capabilities can be invoked by other components without users understanding that images and derived media workflows are being sent off-platform.
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
print(f"Invoking model: {args.model} ...")
try:
response = requests.post(endpoint, headers=headers, json=payload)
response.raise_for_status()
result = response.json()
The tool sends prompts and, when local files are provided, entire base64-encoded file contents to a remote endpoint without any explicit warning, consent step, or data-minimization control. In this skill context, that is risky because users may expect a simple text-to-image helper, not silent transfer of potentially sensitive local media and metadata to an external service.
The primary natural-language description is written only in Chinese and does not indicate that the user can choose another language. Because policy violations include forcing a specific language or locale without user opt-in, this should be documented as optional or multilingual if not region-specific.
No suspicious patterns detected.