Back to skill

Security audit

yby6-video-parser

Security checks for vulnerabilities and agentic risk

Overview

The skill's video parsing and transcription purpose is real, but it handles untrusted URLs and media with weak network and file-retention safeguards.

Review before installing. Use this only in a sandboxed environment for URLs you trust, avoid caller-supplied parse_result data, enable automatic cleanup for sensitive videos, and assume audio files are uploaded to SiliconFlow for transcription while raw media may remain on disk by default.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/transcribe.py:137
Finding

Unrestricted Media Download Enables Server-Side Request Forgery and Resource Exhaustion

Content
View full analysis
{output_path}") headers = {"User-Agent": USER_AGENT} try: response = requests.get(video_url, headers=headers, stream=True, timeout=300) response.raise_for_status() with open(output_path, "wb") as f: for chunk in response.iter_content(chunk_size=8192): if chunk: f.write(chunk) print("视频下载完成") return True except Exception as e: print(f"下载视频失败: {e}") return False ``` The URL reaching this function can come directly from caller-supplied `parse_result` data: ```python # Get the video title to create the temporary directory. data = parse_result.get("data", {}) title = data.get("title", "未命名视频") tmp_dir = create_tmp_dir(title) final_result = { "parse_info": parse_result, "transcription": None } video_url = data.get("video_url") if video_url: temp_id = uuid.uuid4().hex video_file = tmp_dir / f"video_{temp_id}.mp4" audio_file = tmp_dir / f"audio_{temp_id}.mp3" try: if download_video(video_url, video_file): if extract_audio(video_file, audio_file): text = transcribe_audio(audio_file, api_key, model) final_result["transcription"] = text ``` ### Technical Analysis The application performs an outbound GET request to `video_url` without validating: - The URL scheme - The destination hostname - The destination's resolved IP address - Whether the destination is loopback, private, link-local, multicast, or a cloud ...[truncated 2487 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/parser/__init__.py:140
Finding

Substring-Based Platform Detection Allows Requests to Attacker-Controlled Hosts

Content
View full analysis
VideoInfo: """ Parse a sharing link and obtain video information. """ source = "" for item_source, item_source_info in video_source_info_mapping.items(): for item_url_domain in item_source_info["domain_list"]: if item_url_domain in share_url: source = item_source break if source: break ``` For example, the RedBook parser directly requests the URL accepted by this check: ```python async def parse_share_url(self, share_url: str) -> VideoInfo: headers = { "User-Agent": fake_useragent.UserAgent(os=["windows"]).random, } async with httpx.AsyncClient(follow_redirects=True) as client: response = await client.get(share_url, headers=headers) response.raise_for_status() ``` ### Technical Analysis Platform selection is based on whether a trusted domain string appears anywhere in the complete URL. It does not parse and compare the actual hostname. For example, the following attacker-controlled hostname contains the trusted string `www.xiaohongshu.com` and will be classified as RedBook: ```text https://www.xiaohongshu.com.attacker.example/content ``` Because the RedBook parser requests the original URL with redirects enabled, this causes the Skill to contact an attacker-controlled host. The attacker can then redirect the client to another destination. Depending on DNS configuration, DNS rebinding may also cause the accepted hostname to resolve to a private address after validation. Some individual parsers perform their own hostname checks, but this is not consistently implemented. The central dispatch function ther ...[truncated 1536 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/transcribe.py:293
Finding

Sensitive Video and Audio Files Are Retained by Default in Predictable Local Directories

Content
View full analysis
|\n\r\t]', '_', title)[:50] tmp_dir = Path("tmp") / safe_title tmp_dir.mkdir(parents=True, exist_ok=True) return tmp_dir ``` ### Technical Analysis Video transcription only requires media files while downloading, extracting, and submitting audio. Retaining those files after processing is not necessary for the core operation. The default configuration sets `a ...[truncated 1602 chars]
Remediation
View remediation
Vulnerability Patterns
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (83)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill claims broad multi-platform parsing, image-post support, batch use cases, and automatic speech transcription, but the available evidence in the findings indicates the implementation is materially narrower or incomplete. This mismatch is dangerous because users may rely on undocumented or nonexistent safeguards and may unknowingly trigger downloading, parsing, or external transmission behavior under false assumptions about what the skill actually does.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
README.md:29

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
SKILL.md:44