Back to skill

Security audit

Video Analyzer CN

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly matches its video-analysis purpose, but its downloader is too broadly scoped and can fetch arbitrary URLs and overwrite user-writable files.

Install only if you are comfortable with an agent downloading third-party videos, creating local video/frame files, using browser automation for Douyin, and sending frames to a local Ollama service. The downloader should be fixed or constrained before use: allow only intended video domains, write only inside a generated temp directory, avoid overwriting files, stream with size limits, and ask before downloading or processing media.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/douyin_download.py:4
Finding

Unrestricted URL Retrieval and Arbitrary File Overwrite

Content
View full analysis
{output_path}') return output_path if __name__ == '__main__': url = sys.argv[1] output = sys.argv[2] if len(sys.argv) > 2 else r'C:\Users\39535\.openclaw\workspace\tmp\douyin.mp4' download(url, output) ``` The same unsafe, unbounded download pattern is also presented as reusable code in `references/download.md`, lines 24-43. ### Technical Analysis The script accepts both `video_url` and `output_path` directly from command-line arguments. Neither value is validated before use. The URL is passed to `urllib.request.urlopen` without restricting its scheme, hostname, resolved IP address, port, or redirect destination. Consequently, the downloader is not limited to Douyin resources despite its intended purpose. Depending on supported URL handlers and runtime configuration, it may access arbitrary HTTP or HTTPS endpoints, internal network services, loopback services, or local resources through schemes such as `file:`. Redirects are followed without validating the final destination. Therefore, validating only an initial Douyin-looking URL outside this function would not be sufficient: an allowed endpoint could redirect the request to a private or otherwise prohibited destination. The output p ...[truncated 2493 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (14)

Tp2

High
Category
MCP Tool Poisoning
Confidence
85% confidence
Finding

Mixing characters from multiple Unicode scripts in a single identifier is a common technique to create visually ambiguous tool names.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
90% confidence
Finding

The declared description presents an end-to-end video analysis workflow: input a supported platform video URL, automatically download the video, extract frames, analyze them with a local model, and synthesize a summary. The supplied code chunk implements only one intermediate stage: analyzing already-existing local JPG frames by prefix. It does not handle URLs, platform-specific support, downloading, or frame extraction, and its final output is just a count plus individual descriptions rather than a true integrated summary. Calling localhost is not inherently inconsistent with 'no cloud API', since it is still local, but the larger mismatch is that the code covers only a partial subtask rather than the declared full tool behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents a complete multi-platform video analysis pipeline, including downloading, frame extraction, local model inference, and summarization. The actual code chunk implements only a narrow downloader for Douyin video URLs and does not analyze video content, extract frames, invoke any AI model, or summarize results. This is a material description-to-behavior mismatch rather than a mere partial implementation detail.

Content

No source excerpt is available for this finding.

YARA rule 'info_stealer': Information stealer patterns (credential harvesting, browser data theft) [malware]

High
Category
YARA Match
Confidence
83% confidence
Finding

The skill instructs use of Chrome/browser automation to open Douyin pages and discusses browser cookie extraction behavior, which overlaps with credential-access patterns commonly seen in data-stealing tooling. Even though the text says not to use --cookies-from-browser chrome, the workflow still normalizes browser-assisted access to authenticated content and increases the risk of mishandling sensitive browser session data.

Content

Scanner excerpt · SKILL.md (reported line 84)May include surrounding context.

md
`C:\Users\39535\.openclaw\workspace\tmp\`:
- 视频文件: `*.mp4`
- 抽帧图片: `frame_*.jpg` / `dy_frame_*.jpg` / `tt_frame_*.jpg`
- 分析脚本: `analyze_*.py` / `douyin_download*.py`

分析完成后可清理临时文件。

## 注意事项

- 抖音需要Chrome浏览器打开页面提取视频URL(需浏览器MCP)
- 抖音视频URL有时效性,提取后需立即下载
- Chrome v20 cookies加密问题,不能用 `--cookies-from-browser chrome`
- 长视频帧数多时处理时间较长(每帧约20-30秒)
- 内存紧张时优先清理其他进程

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill declares behavior that reads local reference files and accesses external network resources, but it does not declare any explicit tool scope or permissions boundaries. In an agent environment, missing scope makes it easier for the skill to be invoked with broader-than-expected file and network access, increasing the risk of unintended downloads or local file exposure.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger list includes very broad generic terms like “视频”, “分析”, and “video”, which can cause the skill to activate on unrelated user requests. Because this skill performs external downloads and local file creation, accidental invocation raises the chance of unintended network activity and processing of untrusted URLs.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill description does not clearly warn users that it downloads untrusted remote media and stores temporary video frames and scripts on disk. Without explicit notice and consent, users may trigger the skill without understanding the privacy, storage, and content-safety implications of fetching and retaining external media locally.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill provides concrete instructions and code for downloading third-party video content from B站、抖音、今日头条 and storing it locally, but it omits any guidance on copyright compliance, consent, platform terms, or handling potentially sensitive user data embedded in videos. In this skill context, that omission is meaningful because the tool is explicitly designed to fetch, save, and analyze remote media, which increases the chance of unauthorized copying or processing of personal data.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The script base64-encodes each image frame and posts it to an HTTP endpoint, which is a network transmission of potentially sensitive user data. Although the code prints progress, it does not clearly disclose that frame contents are being transmitted to a model service, and this behavior is not explained in a user-facing warning beyond the generic script description.

Content

No source excerpt is available for this finding.

Internal Network Request

Medium
Category
Server-Side Request Forgery
Confidence
70% confidence
Finding

Code issues a request to a loopback, link-local, or private-range host. This can reach internal services not meant to be exposed and is a common SSRF pivot.

Content

Scanner excerpt · scripts/analyze_frames.py (reported line 43)May include surrounding context.

python
'stream': False
        }).encode()
        
        req = urllib.request.Request(
            'http://localhost:11434/api/generate',
            data=data,
            headers={'Content-Type': 'application/json'}

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The natural-language strings and documentation are entirely in Chinese, with no indication that language selection is optional or that the tool is region-specific by design. Under the policy, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The default prompt explicitly instructs the model to respond in Chinese, which imposes a language choice on users without offering an option to select their preferred language. This is a natural-language policy issue because the script defaults to a specific locale rather than making it configurable or opt-in.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

The manifest says the skill uses a local minicpm-v model and does not rely on cloud APIs, which suggests fully local processing. This script actually sends each frame over HTTP to http://localhost:11434/api/generate, meaning it depends on a separate local server API rather than invoking a model directly in-process.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

This code performs a network request to an arbitrary URL and writes the returned data to disk. While it prints a completion message afterward, there is no prior warning, confirmation, or descriptive comment/docstring explaining these side effects to the user beyond the minimal function name.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.