Back to skill

Security audit

guaikei多平台视频转文字

Security checks for vulnerabilities and agentic risk

Overview

This skill appears to provide the promised video-to-text service, but it needs Review because it can fetch and upload user-provided files or URLs while its data-handling scope is not fully constrained.

Install only if you are comfortable sending the selected media, prompt, task ID, and related metadata to GuaiKei and its upload storage provider. Avoid using it on confidential files unless you independently trust the provider's retention and deletion promises, and invoke it only with explicit media paths, URLs, or task IDs.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (18)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the implementation primarily acts as a generic network downloader without clearly surfaced transcription boundaries, users and orchestrators may grant it broader trust than warranted. That mismatch can enable abuse of the skill as a file-fetching primitive, even though the declared purpose suggests a narrower media-processing workflow.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the implementation primarily acts as a generic network downloader without clearly surfaced transcription boundaries, users and orchestrators may grant it broader trust than warranted. That mismatch can enable abuse of the skill as a file-fetching primitive, even though the declared purpose suggests a narrower media-processing workflow.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

If the implementation primarily acts as a generic network downloader without clearly surfaced transcription boundaries, users and orchestrators may grant it broader trust than warranted. That mismatch can enable abuse of the skill as a file-fetching primitive, even though the declared purpose suggests a narrower media-processing workflow.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

If the implementation primarily acts as a generic network downloader without clearly surfaced transcription boundaries, users and orchestrators may grant it broader trust than warranted. That mismatch can enable abuse of the skill as a file-fetching primitive, even though the declared purpose suggests a narrower media-processing workflow.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

If the implementation primarily acts as a generic network downloader without clearly surfaced transcription boundaries, users and orchestrators may grant it broader trust than warranted. That mismatch can enable abuse of the skill as a file-fetching primitive, even though the declared purpose suggests a narrower media-processing workflow.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · README.md (reported line 99)May include surrounding context.

md
4. 同时传入文件路径与任务ID,优先执行 `--id`,忽略 `--file`
5. 无自定义 prompt 时,默认完整转录视频全部文字

| 用户自然语言指令                                         | 自动生成命令                                                                                                |
| -------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 视频提取 https://example.com/video.mp4 中的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 把本地 /path/to/your/video.mp4 改成小红书风格的文案      | `node scripts/video2text/index.js --file "/path/to/your/video.mp4" --prompt "改写成小红书风格的文案"`       |

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill declares use of an environment variable token but does not define an explicit tool/permission scope such as allowed-tools or permissions. In practice, this weakens least-privilege guarantees and makes it harder for a host agent to constrain secret access, especially for a skill that also accepts local file paths and remote URLs.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The trigger description is extremely broad and includes many generic content-processing tasks such as summaries, copywriting, meeting notes, and analysis. In an agent environment, this can cause over-invocation on requests that are not clearly about user-provided video/audio, sending data to an external service unexpectedly.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The natural-language invocation guidance instructs automatic command generation from ambiguous user intents and defaults to reusing previous tasks with --id last. In an agent workflow, this can lead to context confusion, unintended reuse of prior data, or exfiltration of unrelated media-derived content to the wrong request.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · SKILL.md (reported line 128)May include surrounding context.

md
当用户用自然语言下达指令时,按以下映射生成命令,保证识别与执行一致:

| 用户自然语言指令                                     | 生成的命令                                                                                                  |
| ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 提取 https://example.com/video.mp4 里的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 总结这个视频的核心观点 https://example.com/video.mp4 | `node scripts/video2text/index.js --file "https://example.com/video.mp4" --prompt "总结这个视频的核心观点"` |

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

This code transmits user-provided file metadata and later a video URL to backend endpoints via requestApi. Although the docstrings describe the functions technically, there is no visible confirmation prompt or user-facing disclosure here that user content references are being sent over the network.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This code uses Chinese-only natural-language strings for errors, warnings, and generated help output throughout the file. That forces a specific language on users without any opt-in or documented regional constraint, which matches the language/locale policy violation category.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This JavaScript file contains multiple natural-language strings shown to users only in Chinese, such as download status and failure messages. The policy forbids forcing a specific language without user opt-in, and there is no indication here of locale selection or documented region-specific scope.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill emits all user-facing warning and status strings in Chinese, including promotional and operational messages, with no indication that language is configurable or limited to a China-specific context. This creates a natural-language policy issue because the skill effectively forces a specific locale on users without opt-in.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The comment makes a privacy and data-handling promise that is not enforced anywhere in this module: the code uploads user video content but performs no deletion, retention control, or verification that the remote object is removed after processing. In a skill that handles potentially sensitive videos from local files and social-media links, this mismatch can mislead users and integrators into assuming stronger privacy guarantees than actually exist, increasing the risk of unintended retention or compliance violations.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The file's comments and user-visible error/operation strings are written only in Chinese, with no indication that language selection is configurable or chosen by the user. This may violate a language/locale policy if skills must not force a specific language without opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

The user-facing error strings are written in Chinese, including failure, timeout, and token guidance messages, with no indication that the skill is China-specific or that users can opt into a language. This creates a natural-language policy concern because the skill effectively enforces a specific language for operational feedback.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The banner text at L07 is in Chinese, and log timestamps use the environment-default locale via toLocaleString() at L22 without any visible user opt-in or language selection. This can violate language/locale policy when a skill assumes a specific language presentation for all users.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.