T07 · Tool Hijacking and Spoofing
- Location
scripts/run.py:14- Finding
Unverified External Dependency Discovery and Execution
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill has a coherent video-summary purpose, but it runs unverified local helper code and processes remote media with weak boundaries, so users should review it carefully before installing.
Install only if you are comfortable with the skill downloading Xiaohongshu media, running curl/ffmpeg/Whisper locally, and executing a separately installed xiaohongshu-extract script. Prefer running it in a sandbox or dedicated workspace, and verify the helper skill source before use.
scripts/run.py:14Unverified External Dependency Discovery and Execution
SKILL.md:23Untrusted Media Content Is Processed Without a Prompt-Injection Boundary
代码的核心行为是:1)调用外部脚本提取小红书元数据;2)下载视频文件;3)用 ffmpeg 提取音频;4)用 Whisper 生成转录文本。最终输出只有 xhs_meta.json 和 xhs_temp.txt。声明中提到的“深度总结视频内容”在代码中完全没有实现,也没有任何摘要、分析或总结逻辑。对于“提取文案”,代码本身也没有明确抽取视频文案字段,只是生成元数据文件,是否含文案取决于外部脚本,因此不能算作该代码块直接实现的能力。因此描述相比实际行为存在明显夸大,属于描述与行为不一致。
The manifest says the skill can '深度总结视频内容' and is triggered when the user asks to summarize or analyze a Xiaohongshu link. In this file, the implemented behavior stops at metadata extraction, video download, audio extraction, and Whisper transcription; no summarization or analysis step is present, so the actual behavior falls short of the claimed functionality.
The skill invokes shell commands and reads generated files, but it does not declare any explicit tool scope such as permissions or allowed-tools. That creates a governance gap: an agent may execute filesystem and shell operations without clear restriction, increasing the chance of unintended command execution or file access if the skill is triggered in the wrong context.
The trigger language is broad enough that normal requests to analyze or summarize content may activate this skill even when the user did not intend a shell-based video processing workflow. Overbroad activation increases the risk of unnecessary external fetching, local command execution, and handling of untrusted URLs in situations where a simpler, safer response would have sufficed.
The workflow says to run whenever a user provides a Xiaohongshu video link, without clarifying exceptions, consent checks, or validation requirements. In a skill that performs shell execution and downloads remote content, ambiguous activation materially raises the chance of processing attacker-controlled links or invoking the pipeline when the user only mentioned a link contextually.
The skill's implementation performs broad external program execution and chains into another skill's script, increasing the trusted computing base and attack surface. In the context of a link-analysis/summarization skill, this is more dangerous because it combines network retrieval, third-party code execution, and media parsing without clear isolation.
The script dynamically locates and executes another skill's Python script from several filesystem locations, including user-scoped directories. This creates a trust-boundary issue: if that dependency is replaced or tampered with, this skill will execute arbitrary code under the current user context.
# 提取元数据
try:
subprocess.run([sys.executable, extract_script, url, "--flat-only", "--output", "xhs_meta.json"], check=True)
except subprocess.CalledProcessError:
print("Error: Metadata extraction script failed.")
sys.exit(1)
The skill downloads and processes a URL derived from untrusted metadata using an external program. Although arguments are passed as a list and avoid shell injection, this still enables unvalidated outbound network access and retrieval of arbitrary content, which can be abused for SSRF-style access to internal resources or downloading unexpected large/malicious files if the extractor is compromised or returns attacker-controlled URLs.
sys.exit(1)
print("2/4 Downloading video...")
subprocess.run(["curl", "-s", "-L", video_url, "-o", "xhs_temp.mp4"], check=True)
print("3/4 Extracting audio...")
subprocess.run(["ffmpeg", "-i", "xhs_temp.mp4", "-vn", "-acodec", "libmp3lame", "-q:a", "2", "xhs_temp.mp3", "-y", "-loglevel", "error"], check=True)
The skill invokes ffmpeg on a file downloaded from an untrusted remote source. Media parsers have a long history of memory corruption and denial-of-service issues, so automatically feeding attacker-controlled content into ffmpeg expands the attack surface beyond simple summarization.
subprocess.run(["curl", "-s", "-L", video_url, "-o", "xhs_temp.mp4"], check=True)
print("3/4 Extracting audio...")
subprocess.run(["ffmpeg", "-i", "xhs_temp.mp4", "-vn", "-acodec", "libmp3lame", "-q:a", "2", "xhs_temp.mp3", "-y", "-loglevel", "error"], check=True)
print("4/4 Transcribing audio with Whisper...")
subprocess.run(["whisper", "xhs_temp.mp3", "--model", "base", "--language", "zh", "--output_dir", ".", "--output_format", "txt"], check=True)
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
subprocess.run(["ffmpeg", "-i", "xhs_temp.mp4", "-vn", "-acodec", "libmp3lame", "-q:a", "2", "xhs_temp.mp3", "-y", "-loglevel", "error"], check=True)
print("4/4 Transcribing audio with Whisper...")
subprocess.run(["whisper", "xhs_temp.mp3", "--model", "base", "--language", "zh", "--output_dir", ".", "--output_format", "txt"], check=True)
print("======================================")
print("Done! Outputs generated:")
The Whisper invocation hard-codes --language zh, which imposes a specific language/locale behavior on all users. There is no user opt-in, configuration path, or visible justification in this file for restricting transcription to Chinese only.
No suspicious patterns detected.