Back to skill

Security audit

视频配音生成

Security checks for vulnerabilities and agentic risk

Overview

This skill mostly matches a Chinese voiceover workflow, but it also includes a less clearly scoped dubbing path that transcribes video audio, clones a speaker voice, and renders a dubbed video.

Review before installing. Use this only for workflows where Chinese narration or Chinese dubbing is intended, and only with media and voices you are authorized to process. MiMo/Fish/IndexTTS modes may transmit narration text, source audio, or reference voice material to external services. Keep API keys scoped, prefer a contained work directory, and avoid the dub/voiceclone path unless the speaker consent and output-use rights are clear.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (29)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The documented purpose is segmented narration TTS, but the described behavior expands into video/audio extraction, ASR transcription, translation artifact generation, voice cloning, and final video rendering. This mismatch is dangerous because reviewers and users may grant trust or invoke the skill expecting limited audio synthesis, while it actually processes more sensitive inputs, makes broader external requests, and performs materially different actions than disclosed.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The file’s documented behavior is materially different from the skill’s declared purpose: it performs full English→Chinese dubbing with ASR, voice cloning, and muxing rather than simple narration synthesis from a provided narration.json. This kind of capability drift is dangerous because orchestrators, reviewers, or policy controls may invoke or permit the skill under false assumptions, exposing users to undeclared media transformation and biometric voice-cloning behavior.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The header explicitly states that this code is not the ordinary voiceover path and instead performs dubbing, directly contradicting the enclosing skill description. An explicit internal acknowledgment of capability mismatch is a strong indicator that the skill package is mislabeled, making policy review and runtime authorization less reliable.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The code performs speaker voice cloning using a reference audio clip from the source video, which is a sensitive biometric/synthetic-media capability not disclosed by the enclosing narration skill. Undeclared voice cloning can violate consent, policy, or legal requirements and substantially raises misuse risk because it imitates an identifiable speaker rather than generating generic narration.

Content

No source excerpt is available for this finding.

Direct flow: pathlib.Path.read_bytes (file read) → pathlib.Path.write_bytes (file write)

High
Category
Data Flow
Confidence
80% confidence
Finding

Data flows directly from a source (env vars, files, network) to a sink (network output, exec, file write) without intermediate validation.

Content

Scanner excerpt · scripts/tts_audio.py (reported line 46)May include surrounding context.

python
if not samples:
        # A silent block has no RMS to normalize toward; copy it through with neutral
        # metadata rather than dividing by a zero sample count.
        output_wav.write_bytes(input_wav.read_bytes())
        return {
            "rms_dbfs_before": None,
            "rms_dbfs_after": None,

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill advertises operational behavior that uses environment variables, filesystem access, network calls, and shell execution, but it does not declare any explicit tool scope or permissions boundary. This increases the risk of over-privileged execution, makes review harder, and can allow the skill to access or modify resources beyond what a voiceover/TTS task should need.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The description states the skill synthesizes narration into “中文解说音频,” which imposes a specific language/locale behavior in natural language. The file does not indicate user opt-in, selectable output language, or a documented region-specific reason for this restriction.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The trigger list includes broad phrases such as general TTS and voiceover terms that may match ordinary user requests outside the intended workflow. Over-broad activation can cause the skill to run in contexts where the user did not intend media processing or external TTS transmission, increasing the chance of unintended file access, network use, or disclosure of narration content to third-party providers.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The exception message is presented entirely in Chinese, and similar user-facing strings later in the file are also Chinese-only. Because the file does not offer a language choice, fallback, or documented region-specific justification, this can force a specific locale on users and matches the stated natural-language policy violation criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This raised error message is user-facing natural language and is written only in Chinese. The file contains no indication that users can choose their language or that the skill is explicitly restricted to a Chinese-language context, so this is a locale-policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The file-level description states this workflow replaces English speech with a Chinese translation, and later constants and prompts enforce Chinese output. Because the skill forces a specific language/locale behavior without user opt-in in this file, it matches the natural-language policy violation criteria.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill extracts audio and performs ASR on source media even though the declared input model says the workflow should start from an existing timestamped narration.json. This expands access to source content and causes unexpected transcription of potentially sensitive speech, creating privacy and policy exposure beyond what callers would reasonably expect from a narration-only TTS skill.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The produced artifacts do not match the declared contract of the skill, which means downstream components may trust this skill to emit benign TTS segment metadata while it instead generates transcripts, lint reports, full dubbed audio, and a muxed MP4. Contract mismatch increases the chance of unsafe composition, data leakage into unexpected files, and bypass of workflow-specific validation or retention controls.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/lib.py (reported line 16)May include surrounding context.

python
# ── 配置 ──────────────────────────────────────────────────────────────
DEFAULT_MIMO_API_URL = "https://api.xiaomimimo.com/v1"
DEFAULT_MIMO_TOKEN_PLAN_CLUSTER = "cn"
MIMO_TOKEN_PLAN_API_URLS = {
    "cn": "https://token-plan-cn.xiaomimimo.com/v1",

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest describes a skill that synthesizes narrated audio from timestamped narration.json using TTS providers, but this file also defines MiMo ASR configuration and a dedicated speech-recognition API wrapper. Speech-to-text is a distinct capability not mentioned in the skill description and is not an obvious requirement for generating voiceover from prepared narration text.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This module provisions an ASR model and later provides a reusable API client for speech recognition, even though the stated purpose is to turn narration text into spoken audio. Implementing ASR broadens the skill's capability surface in a way that is not justified by the declared TTS/voiceover workflow.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/lib.py (reported line 25)May include surrounding context.

python
}
DEFAULT_MIMO_ASR_MODEL = "mimo-v2.5-asr"  # speech-to-text
DEFAULT_MIMO_TTS_MODEL = "mimo-v2.5-tts"  # text-to-speech
DEFAULT_FISH_TTS_API_URL = "https://api.fish.audio/v1/tts"
DEFAULT_FISH_TTS_MODEL = "s2.1-pro-free"
DEFAULT_FISH_TTS_REFERENCE_ID = "5653cea4ac83480aaf2bf45406556185"

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The configuration sets a fixed Chinese TTS voice and a Chinese-language style prompt, which constrains output locale by default. The file does not offer a user language choice or document that this skill is intentionally region- or language-specific, so this is a natural-language locale policy concern.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/lib.py (reported line 177)May include surrounding context.

python
str(part) if len(str(part)) <= 240 else str(part)[:237] + "..." for part in cmd
    )
    log(f"运行: {display}")
    return subprocess.run(cmd, capture_output=True, text=True, **kwargs)


def get_video_duration(video_path):

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This function serializes the provided text into a request body and sends it to a remote Fish Audio API, which may transmit user content off-system. While the module docstring describes the transport role, there is no user-facing warning, confirmation, or explicit disclosure in this file about sending narration text to an external provider.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The code hardcodes user-visible exception messages in Chinese for configuration validation, and similar Chinese messages continue throughout the module. This imposes a specific language on users without offering a locale choice or documenting that the skill is intentionally Chinese-only.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

This function sends the provided text to a remote HTTP/HTTPS endpoint as JSON, which may transmit user content off-process or off-host. While the module/docstrings describe transport behavior, there is no explicit user-facing disclosure, confirmation, or warning here that input text will be sent to the configured service.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This file includes natural-language error text in Chinese, which forces a specific language in user-visible output. The policy allows locale constraints only when they are clearly documented and justified or when users are given a choice, neither of which is evident in this file.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

Several returned messages intended for users are written only in Chinese, including rejection explanations and hints. Without an explicit language choice or a documented Chinese-only scope, this is a natural-language policy violation under the locale rule.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The MiMo style instruction strings are hard-coded in Chinese and direct the model's speaking style in that language, with no opt-in or alternative locale handling visible in this file. This creates a language/locale policy issue because the skill imposes a specific language context rather than letting the user choose or clearly constraining the skill to a Chinese-only use case.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.