Back to skill

Security audit

Video Learner

Security checks for vulnerabilities and agentic risk

Overview

This skill openly creates new persistent skills from video content, but its safeguards for untrusted transcripts and user review are too thin for that authority.

Review the complete generated SKILL.md before allowing installation, especially any tools, paths, network use, credentials, memory, or shell instructions it requests. Prefer staging generated skills outside the active skills directory until they have been independently checked.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:28
Finding

Untrusted Video Content Can Influence Persistent Skill Instructions

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 28–38
Vulnerability Type: Persistent instruction poisoning through untrusted content
Risk Level: High

Vulnerable Content

markdown
## Processing Flow

1. Create temp directory in `/tmp/` for video download
2. Download video using yt-dlp or douyin-download
3. Extract audio using ffmpeg
4. Transcribe audio to text using Whisper (local)
5. Analyze text content using the agent's LLM capability
6. Display analysis results to user
7. After user confirmation, generate SKILL.md to ~/.openclaw/workspace/skills/<new-skill-name>/
8. Delete temp video files after processing

Technical Analysis

The documented workflow processes video transcripts as untrusted LLM input and subsequently uses the analysis to generate a callable SKILL.md in the persistent ~/.openclaw/workspace/skills/ directory. A video transcript may contain spoken, displayed, or encoded prompt-injection instructions designed to influence the generated Skill.

The workflow does not specify any trust-boundary separation between transcript content and Agent instructions. It also does not require prompt-injection screening, generated-directive validation, capability restrictions, or an independent security review before installation.

User confirmation reduces exploitability but is not a complete security boundary. The process does not explicitly require showing the complete generated file, its requested capabilities, or its security implications before confirmation. Consequently, attacker-controlled transcript content could be transformed into persistent instructions that affect future Agent invocations.

Attack Path

  1. An attacker publishes a supported video containing instructions intended for an AI Agent.
  2. A user supplies the video URL, invoking this Skill.
  3. The Skill downloads the video, extracts its audio, and transcribes the attacker-controlled content. 4 ...[truncated 1337 chars]
Remediation
View remediation

Remediation Suggestions

  1. Treat every transcript and all extracted video metadata as untrusted quoted data. Explicitly instruct the Agent never to follow commands found in that content.
  2. Add prompt-injection detection for text that addresses the Agent, requests policy changes, asks for tool use, or attempts to define persistent rules.
  3. Generate candidate Skills in a non-loadable staging directory rather than writing directly into ~/.openclaw/workspace/skills/.
  4. Require a deterministic validator to reject generated Skills containing undeclared tool use, network access, credential access, shell execution, persistence behavior, or writes outside approved paths.
  5. Present the complete generated SKILL.md, requested tools, dependencies, write paths, and security warnings before requesting approval.
  6. Require separate confirmation for each sensitive capability rather than treating general installation approval as authorization for all generated behavior.
  7. Perform an independent security review after generation and before moving the Skill into the active skills directory.
  8. Apply least privilege to generated Skills and deny access to credentials, memory, networks, and command-execution tools unless separately justified and authorized.
  9. Preserve provenance metadata identifying the source video and the generated sections so reviewers can trace untrusted content.
  10. Consider generating a non-executable analysis report by default and allowing Skill generation only through an explicit, separate user request.
Vulnerability Patterns
  • Rogue AgentSelf-Modification, Session Persistence
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (1)

Session Persistence

Medium
Category
Rogue Agent
Confidence
93% confidence
Finding

The skill explicitly writes generated content into the persistent skills directory under ~/.openclaw/workspace/skills/, which creates session persistence by installing new callable skills that can survive beyond the current run. In context, this is more dangerous than ordinary file output because the written artifact is executable agent configuration/prompt content derived from untrusted video content and LLM analysis, increasing the risk of persistent prompt injection or unauthorized capability extension if review is weak or bypassed.

Content

Scanner excerpt · SKILL.md (reported line 24)May include surrounding context.

md
This skill requires the following tools and access:
- **Tools**: python3, yt-dlp, ffmpeg, whisper (Python package), node
- **Dependency**: douyin-download (requires node)
- **Write path**: ~/.openclaw/workspace/skills/

## Trigger

Static analysis

No suspicious patterns detected.