Back to skill

Security audit

image-to-video-gen

Security checks for vulnerabilities and agentic risk

Overview

The skill generally does what it claims, but it uploads user images and prompts to Google and handles the Google API key in a way that could expose it.

Review before installing. Use only with images and prompts you are allowed to send to Google, avoid confidential or regulated media, and use a dedicated tightly restricted Google API key. The video download step should be hardened to validate the returned URI host and redirects before attaching the key.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
SKILL.md:154
Finding

API Credential Exfiltration Through an Unvalidated Download URI

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:112
Finding

Google API Key Embedded in Request URLs

Content
View full analysis
1 else 2) poll_response = requests.get(POLL_URL, timeout=10) ``` The detailed-workflow examples repeat the same pattern at lines 282-305. HTTPS protects the request in transit from ordinary passive network observers, but secrets in URLs can be exposed through application diagnostics, exception reporting, HTTP client debugging, proxy or gateway logs, monitoring systems, shell history if examples are adapted, and other URL-recording infrastructure. Polling also causes the credential-bearing URL to be used repeatedly. This network access is necessary for the skill's declared image-analysis and video-generation functionality. However, placing a long-lived credential in a URL is not the least-exposure authentication design where header-based authentication or a provider SDK is available. ### Attack Path 1. A user invokes the skill with a valid `GOOGLE_API_KEY`. 2. The skill constructs Veo and polling URLs containing the raw key in the query string. 3. A local debug facility, HTTP proxy, gateway, telemetry collector, exception handler, or diagnostic log records the full URL. 4. An attacker or lower-privileged operator obtains access to that log or telemetry record. 5. The attacker extracts the `key` query parameter. 6. The att ...[truncated 806 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Rogue AgentSelf-Modification, Session Persistence
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (8)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill sends user-provided images and derived prompt content to external Google APIs, but the documentation does not clearly warn users that local content will leave the machine and be transmitted to third-party cloud services. This creates privacy and compliance risk, especially if users process sensitive, copyrighted, or internal images without realizing they are being uploaded.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest documentation says the skill accepts an image as a local path or URL, implying broader input handling. In the actual runnable example, the script only checks for a timestamped local file under ~/.openclaw/workspace/tibetanProc/ and never downloads or parses a user-supplied URL, so the described behavior does not match the implementation.

Content

No source excerpt is available for this finding.

Cloud Storage Exfiltration

Medium
Category
Data Exfiltration
Confidence
91% confidence
Finding

The call to upload_file sends a local image to cloud-hosted Google infrastructure for analysis. That is expected behavior for this skill, but it is still cloud exfiltration of local data and can be dangerous if users assume processing is local or if the image contains sensitive information.

Content

Scanner excerpt · SKILL.md (reported line 75)May include surrounding context.

md
# Step 2: Analyze with Gemini
genai.configure(api_key=GOOGLE_API_KEY)
model = genai.GenerativeModel("gemini-2.5-flash")
image_file = genai.upload_file(str(IMAGE_PATH), mime_type="image/jpeg")

prompt = """Analyze for cinematic video: describe subject, setting, lighting, textures, 
suggested camera movements (dolly, pan, orbit, zoom, rack focus)."""

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 88)May include surrounding context.

md
f.write(analysis)
print(f"✓ Analysis: {prompt_path.name}")

# Step 3: Create enhanced prompt
enhanced = f"""VIDEO GENERATION PROMPT
Duration: 5 seconds
Quality: High Definition

External Transmission

Medium
Category
Data Exfiltration
Confidence
87% confidence
Finding

This code transmits generated prompts and image-derived content to an external Google API endpoint. In context, that is the intended function of the skill, but it is still a real data egress point that can expose sensitive user content if used on private images without clear consent or policy controls.

Content

Scanner excerpt · SKILL.md (reported line 126)May include surrounding context.

md
}

print("\n🎬 Calling Veo API...")
response = requests.post(VEO_URL, json=payload, timeout=60)

if response.status_code not in [200, 202]:
    print(f"✗ API error: {response.json()}")

Cloud Storage Exfiltration

Medium
Category
Data Exfiltration
Confidence
91% confidence
Finding

This is a duplicate cloud upload path in the detailed workflow, again sending a local image to an external service. The risk is contextual rather than overtly malicious, but it is still a true data-exposure concern because local media leaves the host environment.

Content

Scanner excerpt · SKILL.md (reported line 207)May include surrounding context.

md
model = genai.GenerativeModel("gemini-2.5-flash")

IMAGE_PATH = Path.home() / ".openclaw" / "workspace" / "tibetanProc" / "2604110411_input_image.jpg"
image_file = genai.upload_file(str(IMAGE_PATH), mime_type="image/jpeg")

analysis_prompt = """Analyze this image for cinematic video generation:
1. Main subject and focal point

External Transmission

Medium
Category
Data Exfiltration
Confidence
87% confidence
Finding

This second API call site is another explicit external transmission point to Google's Veo service. While necessary for functionality, it still represents outbound transfer of potentially sensitive prompt and image data to a third party and therefore should be treated as a real privacy/security concern.

Content

Scanner excerpt · SKILL.md (reported line 284)May include surrounding context.

md
VEO_URL = f"https://generativelanguage.googleapis.com/v1beta/models/veo-3.0-generate-001:predictLongRunning?key={GOOGLE_API_KEY}"

response = requests.post(VEO_URL, json=payload, timeout=60)
result = response.json()

if response.status_code in [200, 202]:

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

The documentation lists 2604110411_veo_init_response.json as an output artifact, which indicates the skill persists the initial Veo API response. However, the code only parses response.json() in memory and does not write any init-response JSON file to disk, so the documented output contradicts actual behavior.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.