Back to skill

Security audit

paper-to-video

Security checks for vulnerabilities and agentic risk

Overview

This skill is a disclosed local workflow for making a silent paper explainer video, with expected media tooling and no hidden persistence or data theft found.

Install this only if you are comfortable with a skill that creates local project files, uses FFmpeg, and may install pinned npm packages into its own runtime after you approve setup. Review generated outputs before publishing, and use a normal project directory so its output and runtime folders stay contained.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Output HandlingUnvalidated Output Injection, Cross-Context Output, Unbounded Output
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (28)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared purpose describes a skill that transforms research papers into a silent, general-audience HyperFrames video explainer. The supplied code does not perform paper ingestion, summarization, audience adaptation, video generation, or HyperFrames orchestration. Instead, it only checks for a #root element and a KaTeX bundle, then renders data-tex attributes as display-mode math HTML. This is a materially different primary purpose, so the description does not accurately represent the code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description says the skill turns a research paper into a silent general-audience HyperFrames video, implying content processing or video/explainer generation. The supplied code does not process papers, generate videos, or create explainers. Instead, it is a setup script that prepares an isolated npm runtime, installs specific JavaScript packages, and copies KaTeX and GSAP assets into a vendor directory. This is a materially different primary purpose: environment bootstrapping/build support rather than end-user paper-to-video functionality. While dependency setup could support the broader skill, this code chunk itself does not match the declared behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description suggests a higher-level content generation workflow: taking a research paper and turning it into a silent, general-audience explainer in a declarative HyperFrames video format. The supplied code does not analyze papers, simplify content, generate narration/explanation, or assemble a HyperFrames production. Its concrete function is much narrower: it accepts an existing video, an existing SRT subtitle file, and a timeline/layout file, then uses FFmpeg to burn those subtitles into the video while stripping audio and validating the result. This is a materially different primary purpose rather than a mere implementation detail, so the description does not accurately represent the code's behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description says the skill transforms a research paper into a silent explainer video using HyperFrames. The supplied code does not perform any paper processing, summarization, audience adaptation, or video generation. Instead, it is a support/diagnostic utility that audits the local environment and runtime dependencies by querying package metadata, probing executables, checking files under a runtime directory, and outputting a readiness report. While such checks could support a video-generation skill, this code chunk’s actual purpose is materially different from the declared end-user functionality, so this is a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared purpose is about transforming research papers into explainer videos, but the supplied code does not process papers, generate explanations, or create video content. Instead, it performs a system-level font lookup by invoking fc-match and returning font metadata. While font handling could be a supporting detail within a video-rendering pipeline, this specific chunk is materially unrelated to the stated end-user purpose and introduces undeclared system resource access through subprocess execution and local font inspection.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description suggests an end-to-end transformation from a research paper into a silent, general-audience HyperFrames video. The supplied code does not perform paper ingestion, summarization, audience adaptation, or video generation. Instead, it is a focused utility that loads an existing timeline, extracts beat-level subtitle fields, warns about caption layout height, and emits an SRT file. While caption generation could be a supporting component of a video workflow, the actual code chunk’s primary purpose is materially narrower and different from the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description suggests a content-generation or transformation skill: converting a research paper into a public-friendly silent HyperFrames video. The supplied code does not perform paper ingestion, summarization, narration planning, or video generation. Instead, it is a development utility script that lints existing scene HTML files for contract compliance. This is a materially different primary purpose and introduces undeclared capabilities around filesystem access, regex-based source inspection, and command-line validation. Therefore the description does not accurately represent the code's actual behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared purpose describes a content-generation skill that converts research papers into silent, general-audience HyperFrames videos. The supplied code does not perform any paper ingestion, summarization, audience adaptation, video generation, or HyperFrames authoring. Instead, it is a narrow CLI helper for resolving a provided clock time against a timeline data file and returning the matching clip/beat as JSON. This is a materially different primary purpose, so the description does not accurately represent the code.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is about transforming a research paper into a silent explainer video, implying content generation or video assembly from paper input. The supplied code does not process research papers, generate explanations, create HyperFrames output, or produce general-audience narration/content. Instead, it performs quality-control analysis on an already encoded MP4: it loads a timeline, samples deterministic frames, extracts PNGs, computes YAVG brightness in the top 78% of frames, writes a JSON report, and returns failure codes for near-black clips. This is a materially different primary purpose and includes undeclared capabilities related to video validation rather than explainer generation.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description suggests a content-generation skill that converts research papers into silent, general-audience HyperFrames videos. The supplied code does not process research papers, generate explainers, or create video content. Instead, it is a narrow review utility: it loads a timeline, calculates clip/beat/boundary-based snapshot timestamps, optionally finds two-line caption midpoints, and prints them. This is a materially different primary purpose, so the description does not accurately represent the code's behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description says the skill turns a research paper into a silent general-audience HyperFrames video. The supplied code does not ingest or transform a research paper, generate narration/explanations, or build a HyperFrames video. Instead, it verifies an existing video file against duration/media constraints, checks whether subtitles are burned in or absent, validates SRT content against an authoritative timeline, and can atomically replace a final output file if checks pass. These are materially different capabilities and indicate a verification/publishing utility rather than the declared paper-to-video generation skill.

Content

No source excerpt is available for this finding.

Unvalidated Output Injection

High
Category
Output Handling
Confidence
95% confidence
Finding

Model output is used without validation or sanitization. Unvalidated output injected into downstream contexts (SQL, shell, HTML) enables injection attacks and arbitrary code execution.

Content

Scanner excerpt · scripts/review_encoded_video.py (reported line 56)May include surrounding context.

python
def extract_png(ffmpeg: Path, video: Path, seconds: float, output: Path) -> None:
    subprocess.run(
        [
            str(ffmpeg),
            "-y",

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill declares shell, filesystem, and environment-dependent behavior but does not define any explicit tool scope such as allowed tools or permissions. That omission weakens least-privilege controls and can permit broader command execution or file access than users would reasonably infer from the metadata.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 24)May include surrounding context.

md
## Required interaction

Before Phase 2, obtain the video-text language if the user has not already provided it. Do not ask the user to choose subtitles or duration unless they state an additional requirement:

1. Generate subtitles by default because the video has no audio. Use the video-text language for subtitles unless the user requests another subtitle language or explicitly declines subtitles. Cue length is the agent's choice from the language suggestions in [references/video-script.md](references/video-script.md).

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The manifest-like YAML sets language: zh and caption language: zho, which forces a specific language/locale. Under the policy, locale restrictions should either be user-selectable or clearly justified as region-specific; this file provides neither.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The natural-language examples in the user-facing edit protocol are written exclusively in Chinese, which can impose a specific language expectation on users. The file does not indicate that other languages are supported or that Chinese is optional, so this is a locale/language policy concern.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/bootstrap_runtime.py (reported line 20)May include surrounding context.

python
def run(command, cwd):
    subprocess.run(command, cwd=cwd, check=True)


def replace_directory(source: Path, destination: Path) -> Path:

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/burn_subtitles.py (reported line 105)May include surrounding context.

python
f"PlayResX={width},PlayResY={height}'"
    )
    try:
        subprocess.run(
            [
                str(ffmpeg),
                "-y",

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/check_environment.py (reported line 29)May include surrounding context.

python
path = command if Path(command).exists() else shutil.which(command)
    if not path:
        return {"available": False, "path": None, "version": None}
    result = subprocess.run(
        [str(path), *args],
        capture_output=True,
        text=True,

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/ffmpeg_tools.py (reported line 33)May include surrounding context.

python
def supports_subtitles(binary: Path) -> bool:
    result = subprocess.run(
        [str(binary), "-hide_banner", "-filters"],
        capture_output=True,
        text=True,

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/fonts.py (reported line 18)May include surrounding context.

python
"file": None,
            "index": 0,
        }
    result = subprocess.run(
        [matcher, "-f", "%{family}\n%{file}\n%{index}", font],
        capture_output=True,
        text=True,

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/media_validation.py (reported line 11)May include surrounding context.

python
def probe_media(video: Path) -> dict:
    result = subprocess.run(
        [
            "ffprobe",
            "-v",

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/review_encoded_video.py (reported line 46)May include surrounding context.

python
"null",
        "-",
    ]
    result = subprocess.run(command, text=True, capture_output=True, check=False)
    if result.returncode:
        raise RuntimeError(result.stderr.strip() or f"ffmpeg failed at {seconds}s")
    match = re.search(r"lavfi\.signalstats\.YAVG=([0-9.]+)", result.stdout)

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/review_encoded_video.py (reported line 56)May include surrounding context.

python
def extract_png(ffmpeg: Path, video: Path, seconds: float, output: Path) -> None:
    subprocess.run(
        [
            str(ffmpeg),
            "-y",

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

This code enforces a specific locale-format policy for captions by requiring captions.language to match only a lowercase three-letter ISO 639 code. The file does not show any user opt-in, broader locale handling, or justification for why only this language-code format is permitted, which can conflict with language/locale policy requirements.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.