Back to skill

Security audit

AI视频剪辑Skill

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent local video-editing skill, but it needs review because it broadly auto-processes private media and has unsafe install and path-handling practices.

Review this before installing. Use it only on media directories you explicitly choose, avoid sensitive recordings unless you want speech transcribed, set a dedicated output folder, and do not run it as administrator/root. Prefer a virtual environment with pinned dependencies and remove the ambiguous `whisper` package install instruction. Treat analysis JSON and media filenames as untrusted until path validation and FFmpeg concat escaping are fixed.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
INSTALL.md:14
Finding

Ambiguous and Unpinned Whisper Dependency Installation

Content
View full analysis

Vulnerability Details

File Location: INSTALL.md:14 and INSTALL.md:49
Vulnerability Type: Dependency confusion and unsafe package installation
Risk Level: Medium

Vulnerable Code

bash
pip install whisper openai-whisper moviepy Pillow numpy

The guide later repeats the unnecessary package installation:

bash
pip install whisper
python -c "import whisper; whisper.load_model('base')"

Technical Analysis

The project uses the module supplied by openai-whisper, but the installation guide also tells users to install a separate package named whisper. Installing an ambiguous, unnecessary package by name increases dependency-confusion and package-substitution risk.

The direct installation command also leaves all dependencies unpinned. Package content can therefore change between installations, and resolution depends on the package index configured in the user's environment. Python packages may execute build backend or installation logic during installation, so dependency installation is a code-execution boundary rather than a passive download.

The audit did not establish that the current whisper package is malicious. The vulnerability is the unnecessary and ambiguous dependency instruction, combined with unrestricted package resolution.

Attack Path

  1. A user follows the installation instructions and runs pip install whisper.
  2. Pip resolves the package from the user's configured public or private package index.
  3. An attacker who controls, substitutes, or shadows the resolved package supplies a malicious distribution.
  4. Pip executes the package's build or installation process, or the malicious package executes when imported.
  5. The payload runs with the privileges of the user performing the installation.

Impact Assessment

Successful exploitation can result in arbitrary code execution under the installing user's account. The payload could access files, environme ...[truncated 291 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove every instruction that installs the separate whisper package.

  2. Document only the package actually required by the code:

    bash
    python -m pip install openai-whisper
    
  3. Pin all production dependencies to reviewed versions instead of using unrestricted lower bounds.

  4. Generate a lock file with cryptographic hashes and require hash verification during installation, such as pip install --require-hashes.

  5. Use an approved package index and explicitly configure trusted internal mirrors where applicable.

  6. Run dependency installation in an isolated virtual environment without administrator privileges.

  7. Add automated software-composition analysis and periodic dependency review to detect compromised or vulnerable releases.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/auto_clip.py:102
Finding

FFmpeg Concat Directive Injection Through Unescaped Media Paths

Content
View full analysis

Vulnerability Details

File Location: scripts/auto_clip.py:102-105, consumed at scripts/auto_clip.py:117-133
Vulnerability Type: Injection into an FFmpeg concat-demuxer control file
Risk Level: Medium

Vulnerable Code

python
def generate_concat_list(self, output_path: str) -> str:
    concat_file = os.path.join(output_path, "concat_list.txt")
    
    with open(concat_file, "w", encoding="utf-8") as f:
        for seg in self.segments:
            f.write(f"file '{seg['file']}'\n")
            f.write(f"inpoint {seg['start']:.3f}\n")
            f.write(f"outpoint {seg['start'] + seg['duration']:.3f}\n")
    
    return concat_file

The generated file is then passed to FFmpeg with safe-path validation disabled:

python
cmd = [
    self.ffmpeg_path,
    "-f", "concat",
    "-safe", "0",
    "-i", concat_list,
    "-c:v", "libx264",
    "-preset", "medium",
    "-crf", "23",
    "-vf", f"scale={self.config['output_resolution']},fps={self.config['output_fps']}",
    "-c:a", "aac",
    "-b:a", "128k",
    "-af", f"volume={self.config['voice_volume']}",
    "-y",
    output_file
]

Technical Analysis

seg["file"] originates from the analysis JSON and may also reflect an attacker-controlled media filename. It is interpolated directly into the FFmpeg concat-demuxer syntax without escaping or rejecting single quotes, carriage returns, newlines, backslashes, or control characters.

A path containing a quote and newline can terminate the intended file directive and add new concat directives. This is not operating-system shell injection because subprocess.run correctly receives an argument array and does not use shell=True. It is nevertheless parser injection into FFmpeg's concat control language.

The use of -safe 0 further permits otherwise unsafe absolute paths and protocol-like resource names. Depending on the installed FFmpeg build and protoco ...[truncated 1962 chars]

Remediation
View remediation

Remediation Suggestions

  1. Treat analysis JSON and every media path as untrusted input.
  2. Reject paths containing NUL, carriage-return, newline, or other control characters.
  3. Implement escaping that exactly follows FFmpeg concat-demuxer quoting rules; do not rely on surrounding a raw value with single quotes.
  4. Resolve each path with Path.resolve() and verify that it remains beneath an explicitly approved input directory.
  5. Require paths to reference regular files with approved media extensions.
  6. Reject URLs and protocol-prefixed paths unless remote media is an explicit, secured feature.
  7. Avoid -safe 0 where possible. If absolute paths are required, construct the concat file only from application-generated, validated paths and apply an explicit FFmpeg protocol whitelist.
  8. Prefer passing validated inputs directly as FFmpeg arguments and building a filter graph, avoiding a secondary directive-file parser.
  9. Validate the analysis document against a strict schema, including finite numeric ranges for segment start times and durations.
  10. Add regression tests covering quotes, backslashes, CR/LF characters, absolute paths, traversal sequences, and protocol-like paths.
Vulnerability Patterns
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (76)

Tp2

High
Category
MCP Tool Poisoning
Confidence
85% confidence
Finding

Mixing characters from multiple Unicode scripts in a single identifier is a common technique to create visually ambiguous tool names.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

Claiming a complete end-to-end editing pipeline while apparently only handling export/transcoding creates a trust gap that can affect permissioning and user expectations. In agent systems, capability inflation is dangerous because broad triggers and documentation may route unrelated tasks into a component with local file and shell access.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

Claiming a complete end-to-end editing pipeline while apparently only handling export/transcoding creates a trust gap that can affect permissioning and user expectations. In agent systems, capability inflation is dangerous because broad triggers and documentation may route unrelated tasks into a component with local file and shell access.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

Claiming a complete end-to-end editing pipeline while apparently only handling export/transcoding creates a trust gap that can affect permissioning and user expectations. In agent systems, capability inflation is dangerous because broad triggers and documentation may route unrelated tasks into a component with local file and shell access.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

Claiming a complete end-to-end editing pipeline while apparently only handling export/transcoding creates a trust gap that can affect permissioning and user expectations. In agent systems, capability inflation is dangerous because broad triggers and documentation may route unrelated tasks into a component with local file and shell access.

Content

No source excerpt is available for this finding.

Vague Triggers

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The trigger condition covers essentially any video-editing-related request, making accidental or excessive activation likely. Because the skill can analyze local paths, invoke shell commands, and write outputs automatically, overbroad triggering increases the chance of unnecessary data access and unintended execution.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · INSTALL.md (reported line 38)May include surrounding context.

Linux

bash
sudo apt install ffmpeg

3. 安装Whisper (可选,用于字幕生成)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The README explicitly promotes fully automatic ingestion, processing, and export of video content without disclosing that the skill will automatically read media files and write generated outputs to disk. In an agent-skill context, lack of clear user warning and consent around file processing can lead to unintended handling of sensitive media, accidental bulk processing, and silent creation or overwriting of files.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
96% confidence
Finding

The skill instructs use of local Python scripts, shell commands, and reading/writing user-specified filesystem paths, but it declares no explicit tool scope or permission boundaries. In an agent environment, this can lead to overbroad execution authority, unintended access to local files, and unreviewed writes to arbitrary directories.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Generic phrases such as '自动剪辑' or '批量剪辑' overlap with normal conversation and can cause the skill to activate outside the user's intended context. In a file-processing skill, unintended activation is more dangerous than in a read-only helper because it may lead to filesystem access and output generation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Automatic subtitle generation performs speech-to-text on video audio, which is a form of content extraction from potentially sensitive recordings. Failing to disclose this can surprise users, expose private spoken data in text form, and expand the privacy impact beyond simple video editing.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill states that it will automatically store outputs to a preset path but does not prominently warn users beforehand. Silent or assumed file writes can leak content into synced folders, overwrite expected organization, or place sensitive media in locations the user did not intend.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Forcing zh-CN subtitle language without user opt-in can cause inaccurate transcription, unintended language inference, and privacy issues if users did not request speech recognition at all. While lower severity than arbitrary execution, it reflects unsafe defaults for content processing in a multilingual environment.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This YAML file contains comments, labels, and preset names entirely in Chinese, and does not indicate any user-selectable language or locale option for the configuration itself. Under the policy rule, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

All headings and instructional text in the skill file are written in Chinese, with no indication that the language is user-selectable or that the guide is intended only for a Chinese-speaking audience. This can violate language or locale policy when a skill implicitly forces a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

A language or locale policy violation can occur when a skill forces one language without user opt-in. This file presents all instructions, examples, headings, and parameter explanations only in Chinese, with no alternative language option or justification for the restriction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This code file contains natural-language descriptions and user-facing output exclusively in Chinese, beginning with the module docstring and continuing throughout the script. Because the skill forces a specific language for its interface without any user choice or documented locale constraint, it violates the language/locale policy criterion.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/add_effects.py (reported line 225)May include surrounding context.

python
"-y", output_file
        ]
        
        result = subprocess.run(cmd, capture_output=True)
        
        if result.returncode != 0:
            print(f"[特效] 画面优化失败: {result.stderr}")

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The default configuration forces language to zh, which imposes a specific language/locale choice unless the user overrides it. This is a natural-language policy concern because the tool does not default to auto-detection or ask the user to choose a language first.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/add_effects.py (reported line 156)May include surrounding context.

python
"-ar", "16000", "-ac", "1",
            "-y", audio_file
        ]
        subprocess.run(cmd, capture_output=True)
        return audio_file
    
    def transcribe(self, audio_file: str) -> List[Dict]:

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/add_effects.py (reported line 172)May include surrounding context.

python
"-ar", "16000", "-ac", "1",
            "-y", audio_file
        ]
        subprocess.run(cmd, capture_output=True)
        return audio_file
    
    def transcribe(self, audio_file: str) -> List[Dict]:

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/add_effects.py (reported line 188)May include surrounding context.

python
"-ar", "16000", "-ac", "1",
            "-y", audio_file
        ]
        subprocess.run(cmd, capture_output=True)
        return audio_file
    
    def transcribe(self, audio_file: str) -> List[Dict]:

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/add_effects.py (reported line 205)May include surrounding context.

python
"-ar", "16000", "-ac", "1",
            "-y", audio_file
        ]
        subprocess.run(cmd, capture_output=True)
        return audio_file
    
    def transcribe(self, audio_file: str) -> List[Dict]:

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/add_subtitles.py (reported line 49)May include surrounding context.

python
"-ar", "16000", "-ac", "1",
            "-y", audio_file
        ]
        subprocess.run(cmd, capture_output=True)
        return audio_file
    
    def transcribe(self, audio_file: str) -> List[Dict]:

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/add_subtitles.py (reported line 167)May include surrounding context.

python
"-ar", "16000", "-ac", "1",
            "-y", audio_file
        ]
        subprocess.run(cmd, capture_output=True)
        return audio_file
    
    def transcribe(self, audio_file: str) -> List[Dict]:

Static analysis

No suspicious patterns detected.