Back to skill

Security audit

视频内容理解

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent video-analysis pipeline whose main risk is the expected upload of video/audio-derived data to the configured MiMo API provider.

Install only if you are comfortable sending the video, extracted audio, frames, prompts, and relevant background context to the configured MiMo-compatible API endpoint. Use a dedicated work_dir because the skill creates, overwrites, and deletes generated artifacts there, and expect Chinese-oriented prompts and outputs unless you modify the configuration.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (37)

Tainted flow: 'req' from os.environ.get (line 405, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/lib.py (reported line 406)May include surrounding context.

python
for attempt in range(max_retries):
        try:
            req = urllib.request.Request(endpoint, data=data, headers=headers)
            with urllib.request.urlopen(req, timeout=300) as resp:
                return json.loads(resp.read().decode("utf-8"))
        except urllib.error.HTTPError as e:
            body = _sanitize_api_error(e.read().decode("utf-8", errors="replace"))

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的是一个端到端的视频理解与索引技能,覆盖多个分析子系统和多个输出产物。给定代码片段的实际功能明显更窄,只围绕 ASR 结果的时间/来源证据进行写入、验证和摘要:检查文件 identity(size/mtime_ns)、读取背景 research 中的人名术语表、比较 glossary 修正前后文本、校验 audio.wav 及其 meta 文件、验证 asr_result.json 与 sidecar 的一致性。虽然这可以被视为视频分析流水线中的一个支持性组件,但从该代码片段本身看,它并没有执行所声明的大部分核心能力,且其主要目的更接近“ASR 结果可追溯性与一致性验证”。因此,代码行为与声明用途存在明显不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

这段代码的功能明显比声明更窄,且职责不同。声明描述的是“视频理解”主流水线,直接从视频产出 scenes.json、asr_result.json、vlm_analysis.json、silence_periods.json、timeline_fusion.json、agent_narration_brief.md 等完整分析结果;而实际代码是 consolidate.py,只在已有中间结果基础上做二次整理。它读取 asr_result.json、vlm_analysis.json、background_research.json,调用模型完成两类可选操作:1) 清洗 ASR 文本并输出 asr_clean.json;2) 汇总逐场景分析与对话为 understanding_index.json 和 understanding_index.md。代码中没有视频解码、场景切分、语音识别、静音检测、时间线融合或 brief 写作逻辑,因此其实际能力与声明的主要目的存在实质性偏差。虽然“用于理解、索引或总结视频,也作为后续创作前的分析阶段”中的“索引/理解”部分与该代码部分相关,但声明将其表述为完整视频分析技能,并列出多项该代码并未实现的核心输出,因此应判定为描述与行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

声明描述的是一个端到端的视频理解与索引技能,但提供的代码片段只是一段纯函数的数据清洗模块,用于修复模型写出的索引中的时间戳与人物条目问题。代码明确注明“Pure functions only: no I/O, no model calls”,这与声明中的视频分析流水线差异显著。虽然该模块可能可作为更大视频分析系统中的辅助步骤,但就该代码片段本身而言,其实际行为远小于且不同于声明的核心能力,因此应判定为描述与行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

声明描述的是一个完整的视频理解/分析技能,核心职责是对视频进行多阶段处理并产出 scenes.json、asr_result.json、vlm_analysis.json、silence_periods.json、timeline_fusion.json、agent_narration_brief.md 等结果。但提供的代码片段只覆盖 brief 重建/收尾部分:加载已有 JSON 产物、可选融合 mimo_video_overview/background_research、评估素材丰富度、生成 agent brief、添加 storyboard 头,并打印状态。代码注释还明确说明“deliberately performs no extraction, ASR, VLM, or external API calls”。因此这段代码的实际行为与声明的主要用途存在明显偏差:它不是完整的视频分析实现,而是依赖现有分析缓存的后处理与 brief 生成模块。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

声明描述的是一个完整的视频理解分析流水线,涵盖多种分析阶段和对应输出文件。而此代码片段的职责明显更窄:围绕 storyboard 生成与缓存(source/edited storyboard)、多源帧集合处理、缓存命中后的路径迁移修正,以及向 agent_narration_brief.md 预置 storyboard 说明。它依赖外部已有 scenes、frames、clip_plan_validated 等产物,但本身不执行所宣称的大部分分析能力。虽然 storyboard/brief 可能属于“后续创作前分析阶段”的辅助环节,但该片段的实际主功能与声明的主要能力集合存在显著差异,因此应判定为描述与代码行为不匹配。

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill declares capabilities that imply shell, filesystem, environment, and network access but does not define any explicit tool scope such as permissions or allowed-tools. In practice, this increases blast radius: if the skill is invoked in a permissive runtime, it may access more resources than necessary, including API keys, local files, and external endpoints.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger phrases are broad enough to match many generic requests about understanding or analyzing video, without clear scope limits or exclusions. In an agent environment, this can cause unintended auto-selection of a powerful skill that performs file operations, shell commands, and network/API calls on user-provided media.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
79% confidence
Finding

The option table sets the default narration style to 纪录片, and the overall skill description and outputs are framed entirely in Chinese without indicating that language is user-selectable. This can amount to a language/locale policy issue when a skill implicitly enforces one language absent user opt-in or explicit regional justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The document is written entirely in Chinese and includes language-specific processing rules such as counting Chinese characters differently from non-CJK text. There is no indication that users can choose another language or that the language restriction is required for a region-specific or compliance-specific purpose.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The code base64-encodes extracted audio and sends it to a remote MiMo ASR API, which can expose spoken content and incidental sensitive information to a third party. In a video-understanding skill this is somewhat expected functionality, but the lack of an explicit user-facing disclosure/consent mechanism in this path means privacy-sensitive data may be transmitted without clear awareness or control.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The request hard-codes the ASR language from configuration rather than letting the user choose or auto-detect with confirmation. This is primarily a quality, bias, and user-control issue: it can cause mis-transcription, loss of meaning, or inaccurate indexing when the audio language differs, especially in multilingual videos.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This file contains multiple human-facing strings in Chinese, such as the alias prefix "别名" and guidance to prefer labels like "男子"/"白发女子", embedded alongside otherwise English instructions. That creates a language/locale constraint in the skill output without any user opt-in or justification that the skill is intentionally Chinese-only.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The function emits a fixed Chinese section heading and guidance text for sentence-entry anchors. In this Python file there is no surrounding natural-language indication that the skill is intentionally region-specific or that users can choose their preferred language, so it appears to enforce a locale unconditionally.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The embedded natural-language prompts explicitly instruct the model to clean Chinese ASR and build the index using Chinese instructions, which imposes a specific language/locale behavior. The file does not offer a user language choice or document that this skill is intentionally limited to a Chinese-only context.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This file consistently specifies user-facing behavior and status messaging in Chinese, beginning with section headers and function docstrings such as the scene-detection description. Under the policy rule, forcing a specific language/locale without offering a user choice is a natural-language policy violation.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/detect.py (reported line 106)May include surrounding context.

python
"rawvideo",
        "-",
    ]
    result = subprocess.run(cmd, capture_output=True)
    if result.returncode != 0 or not result.stdout:
        raise RuntimeError(result.stderr.decode("utf-8", errors="replace")[:300])
    data = result.stdout

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The function deletes all matching frame_*.jpg files in the target frames directory before extraction. Although there is an inline code comment explaining the cleanup rationale, there is no user-facing disclosure, confirmation, or visible warning that existing files will be removed, and this is a destructive operation.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/lib.py (reported line 18)May include surrounding context.

python
# ── 配置 ──────────────────────────────────────────────────────────────

DEFAULT_MIMO_API_URL = "https://api.xiaomimimo.com/v1"
DEFAULT_MIMO_TOKEN_PLAN_CLUSTER = "cn"
MIMO_TOKEN_PLAN_API_URLS = {
    "cn": "https://token-plan-cn.xiaomimimo.com/v1",

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The default mimo_video_prompt explicitly instructs the model in Chinese and requires Chinese analysis output. This enforces a specific language behavior in the skill without offering a user language choice or documenting a justified region-specific constraint.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/lib.py (reported line 219)May include surrounding context.

python
if len(display) > 2000:
            display = display[:1997] + "..."
    log(f"运行: {display}")
    return subprocess.run(cmd, capture_output=True, text=True, **kwargs)

def get_video_duration(video_path):
    """获取视频时长(秒)"""

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/storyboard.py (reported line 68)May include surrounding context.

python
return candidate
        fc_match = shutil.which("fc-match")
        if fc_match:
            result = subprocess.run(
                [fc_match, "-f", "%{file}", "sans"],
                capture_output=True, text=True, timeout=10,
            )

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The docstring states the analysis should avoid labelling everyone as "黑衣男子," and multiple user-facing strings are hardcoded in Chinese. This indicates a language/locale preference embedded in the skill behavior without any visible user choice or opt-in mechanism in this file.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The script emits user-facing status and error messages in Chinese and also sets a Chinese default style value, but there is no visible option to select another language or locale. This creates a language/locale policy concern because users are implicitly forced into a specific language experience without opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The file contains multiple hard-coded user-facing log and heading strings in Chinese, such as cache and storyboard status messages, with no indication that the user can choose another language. This can violate a language/locale policy when the skill is expected to operate in the user's preferred language unless explicitly documented as Chinese-only.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.