T02 · Agent Memory Poisoning
- Location
SKILL.md:28- Finding
Untrusted Video Content Can Influence Persistent Skill Instructions
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 28–38
Vulnerability Type: Persistent instruction poisoning through untrusted content
Risk Level: HighVulnerable Content
markdown ## Processing Flow 1. Create temp directory in `/tmp/` for video download 2. Download video using yt-dlp or douyin-download 3. Extract audio using ffmpeg 4. Transcribe audio to text using Whisper (local) 5. Analyze text content using the agent's LLM capability 6. Display analysis results to user 7. After user confirmation, generate SKILL.md to ~/.openclaw/workspace/skills/<new-skill-name>/ 8. Delete temp video files after processingTechnical Analysis
The documented workflow processes video transcripts as untrusted LLM input and subsequently uses the analysis to generate a callable
SKILL.mdin the persistent~/.openclaw/workspace/skills/directory. A video transcript may contain spoken, displayed, or encoded prompt-injection instructions designed to influence the generated Skill.The workflow does not specify any trust-boundary separation between transcript content and Agent instructions. It also does not require prompt-injection screening, generated-directive validation, capability restrictions, or an independent security review before installation.
User confirmation reduces exploitability but is not a complete security boundary. The process does not explicitly require showing the complete generated file, its requested capabilities, or its security implications before confirmation. Consequently, attacker-controlled transcript content could be transformed into persistent instructions that affect future Agent invocations.
Attack Path
- An attacker publishes a supported video containing instructions intended for an AI Agent.
- A user supplies the video URL, invoking this Skill.
- The Skill downloads the video, extracts its audio, and transcribes the attacker-controlled content. 4 ...[truncated 1337 chars]
- Remediation
View remediation
Remediation Suggestions
- Treat every transcript and all extracted video metadata as untrusted quoted data. Explicitly instruct the Agent never to follow commands found in that content.
- Add prompt-injection detection for text that addresses the Agent, requests policy changes, asks for tool use, or attempts to define persistent rules.
- Generate candidate Skills in a non-loadable staging directory rather than writing directly into
~/.openclaw/workspace/skills/. - Require a deterministic validator to reject generated Skills containing undeclared tool use, network access, credential access, shell execution, persistence behavior, or writes outside approved paths.
- Present the complete generated
SKILL.md, requested tools, dependencies, write paths, and security warnings before requesting approval. - Require separate confirmation for each sensitive capability rather than treating general installation approval as authorization for all generated behavior.
- Perform an independent security review after generation and before moving the Skill into the active skills directory.
- Apply least privilege to generated Skills and deny access to credentials, memory, networks, and command-execution tools unless separately justified and authorized.
- Preserve provenance metadata identifying the source video and the generated sections so reviewers can trace untrusted content.
- Consider generating a non-executable analysis report by default and allowing Skill generation only through an explicit, separate user request.
