Back to skill

Security audit

B站横屏视频转视频号竖屏切片

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly matches a video-processing workflow, but it explicitly tells the agent to disable a host safe-delete protection during downloads.

Install only if you are comfortable running video-download and ffmpeg commands in a dedicated working directory. Do not disable host safe-delete globally; prefer keeping cleanup inside a disposable workspace and manually deleting known temporary files.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Warning
Location
SKILL.md:56
Finding

Explicit Bypass of the Host Safe-Delete Control

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, line 56
Vulnerability Type: Safety-control bypass through Skill instructions
Risk Level: Medium

Vulnerable setting:

text
CODEBUDDY_SAFE_DELETE_ENABLED=0

The surrounding instruction directs the agent to set this environment variable on a child process because the controlled environment's safe-delete hook would otherwise intercept deletion of intermediate streams.

Technical Analysis

The Skill explicitly instructs the agent to disable a host filesystem safety mechanism while running the download and DASH-stream merge workflow. This is not merely a missing safeguard: the documentation identifies the safe-delete hook as an obstacle and provides a setting intended to suppress it.

The environment variable is not scoped by the Skill to a verified list of temporary files. Consequently, every deletion performed by the affected child process may occur without the protection normally supplied by the host hook. The Skill text therefore crosses the trust boundary between untrusted Skill instructions and host-enforced safety constraints.

Although deleting intermediate DASH streams is part of the expected media-processing workflow, bypassing the platform control is broader than the legitimate need to clean up known workspace files.

Attack Path

  1. A user requests the Bilibili download and video-conversion workflow described by the Skill.
  2. The agent loads and follows the instructions in SKILL.md.
  3. During the download and stream-merging stage, the agent sets CODEBUDDY_SAFE_DELETE_ENABLED=0 for the child process.
  4. The child process performs cleanup operations without safe-delete interception.
  5. Any deletion initiated within that process scope is no longer subject to the host protection, including deletion beyond the expected intermediate files if the invoked tool or its inputs cause a broader cleanup operation.

Impact Assessment

The byp ...[truncated 572 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove the instruction to set CODEBUDDY_SAFE_DELETE_ENABLED=0.
  2. Keep the host safe-delete hook enabled for all child processes.
  3. Run downloads and media conversion inside a dedicated, path-restricted workspace.
  4. If intermediate files must be removed, delete only explicitly enumerated files beneath the workspace after canonicalizing and validating their paths.
  5. Require explicit user confirmation before any cleanup operation that affects files outside the dedicated temporary directory.
  6. Prefer a host-approved cleanup API or path-scoped exception rather than disabling the safety mechanism for an entire process.
  7. Verify that symbolic links and path traversal cannot redirect cleanup outside the workspace.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (15)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的是一个面向B站视频到微信视频号竖屏短视频的端到端制作工具,核心能力应包括视频获取、内容分析、片段选择、画面裁切/重排、字幕与封面生成等。而实际代码只是 scripts/bench_transcribe.py,一个用于测试 Whisper 转写模型在本机 CPU/内存条件下是否可用及其速度的辅助脚本。它仅接收模型名、beam、wav 文件和线程数,做中文音频转写并打印结果,不处理视频、不做剪辑、不生成竖屏内容,也不覆盖声明中的大部分流程。因此代码实际行为与声明用途存在明显且实质性的偏差。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的是一个覆盖视频下载到竖屏短视频生成的完整自动化视频加工工具;但提供的代码片段只实现了字幕文本后处理:读取 .srt 文件、替换特定误识别术语、统一标点,并回写文件。它既不处理视频媒体本身,也不执行下载、转写、剪辑、画面裁切、字幕烧录或封面制作。虽然字幕纠错可视为整体流程中的一个辅助环节,但该代码片段的实际功能与声明的主要用途相比过于局部,不能准确代表所述技能的核心能力,因此应判定为描述与行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

声明描述的是一个端到端的视频加工技能,主功能应是把横屏长视频转换为竖屏短视频切片,并覆盖多个处理阶段。实际提供的代码只是一个辅助探测脚本:按时间点抽帧、读取图片像素、统计底部内容出现的行,用于估计内嵌字幕带位置。虽然这可能是视频裁剪/字幕处理流程中的一个支持性步骤,但就该代码块本身而言,其主要行为与声明的整体能力差距很大,无法代表所宣称的核心功能,因此构成明显不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

声明描述的是一个端到端的视频加工技能,核心能力应包括视频下载、内容选段、画面重构为9:16竖屏、字幕烧录和封面生成等。而这段代码只是其中很小的一步:对预先指定的片段进行音频截取和 Whisper 转写,输出字幕文件。虽然“转写”属于声明流程中的支持环节,但代码的实际主功能明显比声明窄得多,且缺失声明中的关键视频处理能力。因此描述不能准确代表该代码块的实际行为,属于明显不匹配。

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill clearly instructs use of shell commands, environment variables, and filesystem reads/writes, but it declares no tool scope or permission boundary. In an agent environment, this increases the blast radius because the runtime may grant broader capabilities than users expect, enabling command execution and local file access without explicit disclosure or restriction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The instruction 请使用简体中文输出 explicitly constrains the generated transcription language/script to Simplified Chinese. This is a natural-language locale policy issue because the file does not present it as a user choice or opt-in, and it is framed as the default behavior to enforce.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The prompt explicitly instructs the model to output in Simplified Chinese, and the transcription call later fixes the language to Chinese. This imposes a specific language/locale behavior without offering the user a choice, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The transcribe call sets language="zh", forcing Chinese-language processing regardless of user preference or input content. Because no opt-in or configurable locale mechanism is provided, this is a natural-language policy concern under the language/locale rule.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The module docstring is written as a directive for generating WeChat video covers in Chinese context, and the script also hardcodes Chinese user-facing text. There is no indication that language/locale is optional or limited by an explicitly documented region-specific requirement, which can violate language-choice policy.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/probe_subs_band.py (reported line 14)May include surrounding context.

python
rows_hit = {}
for k in range(N):
    t = dur * (k + 0.5) / N
    subprocess.run([FFMPEG, "-y", "-ss", f"{t:.1f}", "-i", SRC, "-frames:v", "1",
                    "-q:v", "2", TMP], capture_output=True)
    im = Image.open(TMP).convert("RGB")
    W, H = im.size

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/slice_vertical.py (reported line 74)May include surrounding context.

python
cmd += ["-i", src, "-vf", vf,
            "-c:v", "libx264", "-crf", "23", "-preset", "fast",
            "-c:a", "aac", "-b:a", "128k", "-movflags", "+faststart", out]
    r = subprocess.run(cmd, capture_output=True, text=True, encoding="utf-8", errors="replace")
    if r.returncode != 0:
        print("FFMPEG ERROR:\n", r.stderr[-1500:])
    return r.returncode

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The comments and prompt explicitly instruct the transcription model to output Simplified Chinese, and the transcribe call also fixes the language to Chinese. This is a natural-language locale policy issue because the skill does not offer any user-selectable language or script option.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/transcribe_clips.py (reported line 60)May include surrounding context.

python
wav = f"download/clip{idx}.wav"
        dur = end - start
        if not os.path.exists(wav):
            subprocess.run([FFMPEG, "-y", "-ss", str(start), "-t", str(dur),
                            "-i", SRC_WAV, "-ac", "1", "-ar", "16000",
                            "-c:a", "pcm_s16le", wav], capture_output=True)
        print(f"[{idx}] {title} {start}-{end}s ({dur}s) -> {wav}", flush=True)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The code hard-codes language="zh" for transcription, which imposes a specific language/locale behavior regardless of user preference. The file does not indicate that this is optional, user-selected, or justified as a region-specific tool.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This code writes a full transcript JSON file and a preview text file to disk, which may contain sensitive spoken content from the input audio. The script includes status prints but no warning, prompt, or explanatory comment/docstring disclosing that user audio will be persisted as plaintext transcript artifacts.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.