Back to skill

Security audit

抖音视频转报告

Security checks for vulnerabilities and agentic risk

Overview

The skill has a coherent Douyin video reporting purpose, but its implementation creates high-impact command-injection and network-request risks and does not clearly scope user consent or external processing.

Install only after the publisher fixes the shell command construction, restricts URLs to validated Douyin hosts, adds consent before downloading and sending audio/images to external services, narrows triggers, and aligns the implementation with the promised report and Feishu delivery behavior. Run it only in a sandbox with minimal file and network access if testing is necessary.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
douyin_pipeline.py:13
Finding

Shell Command Injection Through a Remotely Controlled Video URL

Content
View full analysis
{ const v = document.querySelector('video'); if (v && v.src && v.src.includes('douyin')) return v.src; return null; } """) ``` ```python def download_video(url, path): print(f"📥 Downloading video...") code, _, err = run( f'curl -L -o "{path}" --max-time 120 -s ' f'-H "User-Agent: Mozilla/5.0 (iPhone; CPU iPhone OS 17_0 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.0 Mobile/15E148 Safari/604.1" ' f'-H "Referer: https://www.douyin.com/" ' f'"{url}"', timeout=130 ) ``` ### Technical Analysis The `run` helper executes command strings through a system shell by setting `shell=True`. The video URL extracted from the loaded page is inserted directly into the curl command without shell-safe argument handling. The only relevant validation checks whether the full URL contains the substring `douyin`. This is not a hostname or syntax validation step. A page controlled by an attacker can expose a crafted `video.src` containing that substring together with shell metacharacters or command-substitution syntax. Surrounding the URL with double quotes does not prevent all command injection. In common POSIX shells, command substitutions such as `$()` remain active inside double-quoted strings. Consequently, a malicious value reaching this command can cause local commands to run under the privileges of the pipeline process. ### Attack Path 1. An attacker provides a URL leading to a page under the attacker's control. 2. The page creates a `video` element whose `src` includes th ...[truncated 1097 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
douyin_pipeline.py:20
Finding

Server-Side Request Forgery Through Unrestricted Page and Media URLs

Content
View full analysis
{ const v = document.querySelector('video'); if (v && v.src && v.src.includes('douyin')) return v.src; return null; } """) ``` ```python code, _, err = run( f'curl -L -o "{path}" --max-time 120 -s ' f'-H "User-Agent: Mozilla/5.0 (iPhone; CPU iPhone OS 17_0 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.0 Mobile/15E148 Safari/604.1" ' f'-H "Referer: https://www.douyin.com/" ' f'"{url}"', timeout=130 ) ``` ### Technical Analysis The pipeline accepts its initial navigation URL directly from a command-line argument and passes it to Playwright without validating the scheme, hostname, resolved address, or redirect chain. The media URL is also insufficiently validated. Testing whether a URL contains the text `douyin` does not establish that it belongs to Douyin. The text can appear in a path, query parameter, username component, or attacker-controlled hostname. Curl then follows redirects with `-L` without validating each destination. ...[truncated 1586 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
SKILL.md:111
Finding

Unpinned Third-Party Package Installation Instructions

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (15)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The skill description claims a narrowly defined reporting workflow, but the observed behavior includes bypassing anti-bot verification, downloading third-party media, and reportedly performing additional analysis/output modes not clearly disclosed. This mismatch is dangerous because users and reviewers cannot accurately assess what data is collected, how content is processed, or where outputs go, increasing the risk of unauthorized processing and exfiltration.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill openly describes bypassing verification, downloading content, and sending outputs externally, yet it provides no user-facing warning, consent checkpoint, or policy boundary for these high-risk actions. This is especially dangerous here because the workflow touches third-party protected content, local file creation, transcription, and external delivery, creating legal, privacy, and misuse risk.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
98% confidence
Finding

This is a true tool-parameter abuse issue because the generic run() wrapper executes arbitrary shell strings, and the skill constructs those strings with untrusted or insufficiently validated data. In an agent context that processes user-supplied links, this significantly raises the risk of command injection and full compromise of the local runtime or accessible workspace.

Content

Scanner excerpt · douyin_pipeline.py (reported line 14)May include surrounding context.

python
WORKSPACE = "/home/gem/workspace/agent/workspace"

def run(cmd, timeout=60):
    r = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=timeout)
    return r.returncode, r.stdout, r.stderr

async def get_video_info(url):

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill advertises and demonstrates shell-capable behavior such as running curl, ffmpeg, and Python scripts, but it does not declare an explicit permission boundary like allowed-tools or permissions. This weakens governance and review because the runtime capabilities exceed what is formally scoped, making unintended command execution or file access harder to constrain.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The trigger phrases include broad natural-language expressions like '这个视频说了什么' and '总结这个视频', which can match ordinary conversation and invoke the skill unintentionally. In this skill's context, accidental activation is more dangerous because invocation can lead to downloading remote content, running browser automation, and processing media without clear user intent.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

文档写明语音识别为“中文普通话识别”,且示例代码将浏览器 locale 固定为“zh-CN”,但没有说明这是用户可选项,也没有将该技能限定为仅适用于中文/中国区场景。这样的自然语言说明构成了未经用户选择的语言/地区强制。

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The description promises a "完整图文报告" and all user-facing metadata is written only in Chinese, implying a fixed language/locale experience without user opt-in. There is no indication that the skill offers language selection or that the Chinese-only behavior is a documented, justified regional constraint.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The trigger list includes broad everyday phrases such as '这个视频说了什么', '总结这个视频', and '抖音链接', which can match normal conversation rather than an explicit request to invoke this skill. In an automation that performs fetching, downloading, transcription, report generation, and outbound sending, overbroad triggers increase the chance of unintended activation and unreviewed processing of third-party content.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
97% confidence
Finding

The helper executes shell commands with shell=True, and this file interpolates externally influenced values such as URLs and file paths directly into command strings. Because the skill ingests an untrusted Douyin URL and later uses derived values in curl/ffmpeg/CLI invocations, a crafted input containing shell metacharacters could lead to arbitrary command execution in the agent environment.

Content

Scanner excerpt · douyin_pipeline.py (reported line 14)May include surrounding context.

python
WORKSPACE = "/home/gem/workspace/agent/workspace"

def run(cmd, timeout=60):
    r = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=timeout)
    return r.returncode, r.stdout, r.stderr

async def get_video_info(url):

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill sends extracted audio to an external speech-to-text CLI service without explicit notice or consent. Since videos may contain personal, confidential, or copyrighted speech, this creates a real privacy and data-governance risk through unexpected third-party transmission and processing.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill analyzes video frames through an external CLI service without warning that screenshots may be uploaded or remotely processed. Frames can contain faces, on-screen messages, account details, or other sensitive visual data, so silent transfer materially increases privacy and compliance risk.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

技能描述声明会自动完成“炫酷HTML报告 → 飞书发送”,但主流程实际只抓取视频、下载、抽帧、转录、逐帧分析,并将结果写入本地 analysis 文件。代码结尾仅打印“下一步: AI 根据画面分析+语音转录生成完整报告”,表明完整报告生成和发送并未在此实现。

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The browser context is hard-coded to locale='zh-CN', which enforces a specific locale behavior. The file does not offer opt-in or explain that the skill is intentionally limited to a China-specific use case, so this is a natural-language locale policy concern.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
76% confidence
Finding

manifest 描述列出的是抓取、下载、音频提取、语音转文字、生成 HTML 报告、飞书发送;而代码还额外执行 ffmpeg 抽帧和 image-understanding 逐帧视觉分析。虽然这可能有助于生成图文报告,但该关键能力未在 manifest 描述中声明,而代码文档却将其作为核心流程。

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

The transcription command forces --lang zh, requiring Chinese output regardless of user preference. There is no indication in the file that the user can choose the language or that the restriction is explicitly documented as part of a region-specific workflow.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.