Back to skill

Security audit

视频成片组装

Security checks for vulnerabilities and agentic risk

Overview

This skill is a local video assembly tool whose file and FFmpeg operations match its stated media-processing purpose.

Before installing, ensure you are comfortable with a local media tool that can run ffmpeg/ffprobe, create and delete files in the selected work/output directories, and overwrite normal recap output aliases. Use fresh work directories for strict adoption paths and review output paths when exporting JianYing drafts or bundling media.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (85)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

Based on the supplied code chunk, the implementation does not substantiate the declared functionality. The only observable behavior is a docstring indicating an adoption-related module focused on narration/audio-mix binding and strict publishing for adopted-source audio. That is materially different from the declared end-to-end video compositing role involving muxing, subtitle generation, burn-in, and normalization. Since no operative code is shown for the declared pipeline, and the module description points to a different primary purpose, this should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The code is focused on explicit audio mix binding and validation. It reads and validates an audio mix adoption JSON, probes the input video only to derive/verify the picture clock, renders narration placements into stereo WAVs, combines them into a voice bus, mixes that with a pre-prepared bed, applies master gain, validates the final rendered file's audio/video timing, and records a finalized JSON report. This is related to the broader declared final-composition stage, but several core declared capabilities are missing from the supplied chunk: there is no creation of SRT/ASS subtitles, no subtitle burn-in, no direct mux/assembly of video streams, and no implementation of ducking against source audio. The code also does not perform loudness normalization such as EBU R128/ITU loudness processing; it only applies a configured gain value. Because multiple prominently declared end-user capabilities are not represented in the actual code behavior, the description does not accurately match this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的是一个最终成片组装技能,核心能力应围绕 mux、mix、ducking、subtitles 和 loudness normalization。给出的代码却只处理时间基与包级时钟验证:probe_picture 通过探测视频流/帧/包信息验证零起点 CFR 时钟与 MP4/H264/HEVC 约束;validate_aac_packet_interval 检查 AAC 包 duration、PTS/DTS 连续性以及与流头时间区间的一致性;validate_pair_timing 仅判断音视频区间是否在容差内匹配。代码注释还明确说明“not a perceptual synchronization verdict”。因此其主要目的与声明显著不符,属于功能方向上的实质性不匹配,而非仅仅是支持性实现细节。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

声明描述的是一个完整的最终合成阶段技能,核心能力应包括音视频合成、混音、字幕处理和响度处理。但提供的代码文件 frozen_audio.py 的职责非常明确且狭窄:调用 ffprobe 读取音频/视频流与 packet 元数据,验证某个输入音频流是否为可直接拷贝的 AAC、其起止时间是否覆盖主画面区间,并在输出文件中逐项比对 packet 数量、字节数、PTS/DTS、duration 和 side data 是否保持不变。这是媒体探测/校验工具,不是成片合成实现。虽然它可能作为更大视频装配流程中的辅助校验模块存在,但就这段代码本身而言,其实际行为与声明的主要用途存在明显且重大的不一致。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description is for a media-processing skill whose main job is final video composition. This code chunk does not implement that pipeline. Instead, it focuses on strict validation of narration adoption metadata, copying/snapshotting input audio files, probing audio characteristics via ffprobe, tracking derived audio assets, and emitting a JSON binding/provenance record. While these steps could support a larger video assembly workflow, the chunk itself lacks the described core behaviors: no FFmpeg render command for video composition, no subtitle generation, no audio ducking/mixing logic, and no loudness normalization. Therefore the code's actual behavior is materially different from the declared primary purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

声明描述的是媒体处理能力本身:混音、ducking、字幕生成/烧录、响度标准化、输出最终成片与字幕。但给出的代码片段并未执行这些音视频处理步骤,也没有生成字幕、烧录字幕、调用编码器/混音器,或实现响度标准化算法。它主要负责严格发布流程中的状态管理与质量门禁:写入 narration/audio mix 的绑定记录,调用 assembly QC,发布候选输出到正式路径,并在异常时回滚相关文件。虽然这些操作可能属于更大“最终合成阶段”的配套发布环节,但就该代码片段本身而言,其主要目的与声明的核心功能不一致,因此应判定为描述与行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The declared purpose is media composition: combining narration audio with a source video into a final output. The supplied code chunk does not perform any audio/video processing, muxing, rendering, or media manipulation. Instead, it gathers file identity metadata, resolves a configured source video path, checks artifact existence, and reads JSON files such as timeline/provenance data from disk. Those behaviors suggest provenance/cache/status tracking rather than final video synthesis. This is a materially different primary purpose from the declared description, so it should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的是“最终合成”能力:把旁白铺到视频上、做 ducking、生成并可烧录字幕、响度标准化,最终输出 recap 成片。这些都是渲染/后期处理行为。实际代码却仅实现了一个 JianYing/CapCut draft exporter:从 timeline 构建剪映草稿 JSON,调用 ffprobe 获取媒体元信息,并将草稿写入目录,供用户在剪映中打开和继续编辑。虽然注释中提到时间线可包含旁白、BGM、字幕和音量关键帧,但本代码本身并不执行这些媒体处理,也不产出最终视频或字幕文件。因此其主要用途与声明明显不一致,属于实质性描述-行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

There is a clear description-to-code mismatch based on the supplied chunk. The description promises a full final compositing pipeline for recap video production, including narration overlay, source-audio ducking, subtitle creation/burn-in, and loudness normalization. However, the only actual code shown is a module docstring for a Jianying draft export subpackage, with no implementation of any of those declared functions. The visible purpose also points in a different direction (draft export for 剪映) rather than final recap assembly. While this could be only a partial file, based strictly on the supplied code chunk the declared description is not accurately represented.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description says this skill is the final composition stage that takes source video, narration metadata, and narration positions, then mixes audio with ducking, generates subtitle files, optionally burns subtitles into the video, and normalizes loudness to produce a final output video. The supplied code chunk instead is a builder module for a JianYing exporter. It constructs internal timeline/material/segment data for video, audio, text/subtitle, and image tracks, including audio volume keyframes and subtitle-like text segments. While volume keyframes and subtitle track construction are related supporting pieces, the code does not actually assemble/export a final media file, does not create SRT/ASS files, does not burn subtitles, and does not normalize loudness. So the actual code represents project/timeline-building infrastructure, not the declared end-to-end final video synthesis behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description is for an end-stage media synthesis pipeline that should manipulate actual audio/video assets and produce a final recap file plus subtitles. The supplied code chunk is only an internal data model (DraftBuildContext) used to normalize timeline state and manage track allocation/order. Its functions are limited to constructing context from timeline metadata, adding segments to tracks, finalizing track flags/order, and storing notes. These are supporting timeline-authoring utilities, not the described final composition behavior. Because the primary purpose and major capabilities in the description are absent from the code shown, this is a clear description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description centers on final media composition capabilities: video/audio muxing, ducking, subtitle generation/burn-in, and loudness normalization. The actual code chunk does none of those things. Instead, it optionally imports a JianYing exporter, loads timeline.json, and writes a JianYing draft project to a target directory, explicitly as a non-critical sidecar that does not affect the already rendered recap. This is a materially different purpose and adds an undeclared capability (exporting an editing-project draft). Therefore the description does not accurately represent this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的是一个最终成片合成技能,核心能力应包括媒体混音、字幕生成/烧录和音频响度处理。但给出的代码片段是纯数据结构与模板装配代码,服务于剪映导出草稿的 schema/metadata 构建,不涉及媒体文件处理、音视频编解码、字幕生成、时间窗压低原声或响度归一化。虽然这类 schema 代码可能是更大系统的辅助部分,但就该代码片段本身而言,其实际行为与声明的主要用途明显不一致。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

This code chunk is a narrow helper module for reading pinned protocol templates from disk and returning mutable deep copies. Its behavior is limited to file lookup, JSON loading, caching, and copying. That is materially different from the declared primary purpose of final-stage video synthesis and muxing with narration, subtitles, ducking, and loudness processing. While this helper could support a larger video assembly pipeline, the supplied code itself does not carry out the described operations, so the description does not accurately represent this chunk’s actual behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

声明描述的是一个完整的视频最终合成器,涉及媒体处理、字幕生成/烧录、音频压混和响度标准化。但提供的代码片段没有任何媒体处理逻辑,也不读取/写出视频、音频或字幕文件;它只对 timeline 对象进行结构校验,确保字段存在、类型正确、时间区间有效,并拒绝某些不再支持的字段。这不是对声明功能的辅助实现细节,而是一个不同层级、不同目的的数据验证模块。因此描述与该代码片段实际行为存在明显不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

声明描述的是一个完整的“最终合成”技能,核心能力应包括音视频 mux/混音、ducking、字幕生成与烧录、响度处理以及成片输出。但提供的代码片段只是一个辅助性的数据结构与算法模块,用于为不同类型轨道设置固定布局顺序,并在时间区间重叠时为同名轨道分配确定性的后缀名称。这与声明中的主要功能存在明显差异。虽然轨道分配可能是视频导出流程中的支持性实现细节,但就该代码片段本身而言,其行为远不足以支撑所声明的技能目的,因此应判定为描述与代码行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

声明描述的是“最终成片合成”阶段,核心能力应包括媒体时间线合成、音频 ducking、字幕生成/烧录以及响度处理。但提供的代码仅处理 JianYing 草稿文件的安全落盘和素材打包:复制资源文件、计算 MD5、填充 draft meta、避免目录冲突、原子写 JSON。代码没有调用任何 FFmpeg/媒体处理逻辑,也没有字幕格式生成、音频分析、响度处理或视频编码输出。因此其主要用途与声明严重不符。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

声明描述的是“最终合成”核心阶段,重点能力应包括音频混音、旁白驱动的 ducking、字幕生成/烧录以及响度标准化。但提供的代码片段并未执行这些操作。它主要是辅助性的媒体分析与来源映射模块:检查 clip_plan_validated.json 是否过期、构建源片段到时间线的映射、探测视频尺寸/旋转/SAR/DAR/FPS、检查音频流是否存在、读取视频编码与色彩信息,并生成色彩标签和是否允许视频流拷贝的判定。这些可以服务于视频组装流程,但与声明中的主要功能相比差异明显,且当前片段的实际主用途并不是“把旁白和字幕合成到最终视频中”。因此应判定为描述与代码行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

声明描述的是“最终成片合成”核心流程,重点在音频混音、ducking、字幕生成/烧录和响度标准化;而这段代码的实际职责是视频上的静态图片包装层叠加及其元数据导出。虽然“烧录式视觉合成”与最终合成阶段存在弱相关,但该代码片段没有体现声明中的主要能力,反而实现了描述中未提及的 packaging/frame/logo 覆盖功能。因此该代码块与声明用途存在明显的描述—行为不一致。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的是一个最终成片组装器,核心对象应是视频+旁白+字幕的最终 mux/burn/normalize 流程。该代码虽然会探测视频流的 CFR 属性,但目的只是验证 source segment 的帧时钟并据此换算采样点;后续真正执行的都是 ffmpeg 音频解码、裁切、淡入淡出、延时对齐、amix 混合,产出 48kHz 立体声 float WAV 文件(source_bed、score_bed、prepared_bed)和 JSON receipt。代码没有处理旁白轨本身,没有把任何音频 mux 回视频容器,没有字幕文件生成或烧录逻辑,也没有 loudnorm/ebur128 一类响度标准化,仅用 astats 检查峰值和非有限值。因此其主要目的与声明严重不符,应判定为描述与行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

声明描述的是一个最终成片装配/混音模块,其核心能力应包括视频与音频合成、原声压低、字幕生成/烧录和响度规范化。但该代码块仅处理字幕数据层逻辑:加载 original_subtitles、user_subtitles、ASR 结果,解析 SRT/ASS,计算 narration gaps,把原始对白字幕放入旁白空档,并评估 source subtitle mask policy。它不调用 ffmpeg、不处理媒体流、不输出成片,也不执行音频混音或响度标准化。因此,这段代码的实际主用途与声明的主用途明显不一致。虽然它与“字幕”这一大功能域相关,可能是整个装配流程的辅助子模块,但就该代码块本身而言,声明未准确反映其实际行为。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

There is a clear mismatch between the broad declared end-to-end final composition functionality and the supplied code chunk, which is only an empty init.py with a descriptive docstring for subtitles. The code shown does not demonstrate the primary declared behaviors. While the docstring suggests some subtitle-related scope, the implementation excerpt is too limited and does not substantiate the claimed capabilities of assembling a final video, muxing audio, ducking, burning subtitles, or normalizing loudness.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The description claims an end-to-end final compositing capability for video/audio/subtitles. However, this code only operates on subtitle track metadata using the Python standard library. It parses and validates a versioned subtitle-track JSON structure, checks that referenced picture/audio identities match expected values, validates cue timing and evidence, and returns structured metadata plus cue entries. It does not invoke media tools, process video or audio streams, generate subtitle formats like SRT/ASS, burn subtitles into video, or normalize loudness. While subtitle validation could be a supporting component within a larger video-assembly system, this chunk by itself materially differs from the declared primary purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

声明描述的是一个“最终合成”技能,核心应包括音频铺轨、混音、ducking、字幕生成/烧录以及响度标准化。但该代码片段没有执行任何音视频合成、转码、混音或响度处理,也没有生成 recap 成片。它的实际职责是对已有 subtitle_track 做严格的媒体绑定和时钟投影验证,确保字幕 cue 与输入/输出视频的帧时钟一致,并防止文件变更导致的陈旧绑定。这与“assemble video / mux / ducking / 成片”这类声明主用途存在明显偏差。虽然字幕相关与声明有部分交集,但这里只是字幕轨验证与证据管理的子模块,而不是所声明的最终视频合成功能。

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The documentation describes a standalone source/score bed producer and explicitly limits its scope, while the skill metadata claims this skill performs final video assembly with narration, ducking, subtitle generation, burn-in, and loudness normalization. This mismatch is dangerous because downstream agents or operators may trust the skill for end-to-end final assembly and accidentally skip required validation or processing steps, producing incorrect outputs or misrepresenting partially verified audio as a finished deliverable.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.