Back to skill

Security audit

guaikei视频转文字字幕生成

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly does what it claims, but it can fetch arbitrary URLs and upload media to remote services while its documentation understates that network scope.

Install only if you are comfortable sending selected videos, downloaded URL content, prompts, and a GUAIKEI_API_TOKEN-authenticated request to the provider. Avoid using it on confidential meetings, proprietary recordings, internal URLs, localhost/private-network URLs, or regulated data unless you have separately verified the provider's retention and privacy claims.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (23)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

声明描述的是一个“视频转文字/总结”技能,核心应涉及视频输入处理、媒体解析、语音识别或文本生成等能力。但提供的代码片段只是一段通用参数解析器(parseArgs、buildHelp),用于处理命令行选项、校验参数、生成帮助文本,不包含任何视频、链接、字幕、转写或总结相关逻辑。这不是单纯的底层支持细节差异,而是代码行为与声明主功能完全不一致,因此应判定为明显不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

代码内容与“视频转文字/字幕提取/总结”这一声明用途不匹配。该代码没有看到任何音视频解析、语音识别、字幕抽取、文本生成或总结逻辑;相反,它实现的是一个通用下载器组件,负责通过 HTTP/HTTPS 请求获取远程资源并保存到本地文件系统。虽然下载公开视频链接中的视频可能是视频转写流程的辅助步骤,但当前代码块本身只体现了下载能力,而且是可用于任意文件的通用能力,属于未在描述中明确声明的实际行为。因此应判定为描述与代码行为存在实质性不一致。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

声明描述的是一个视频转文字/视频内容提取类技能,但提供的代码片段只实现了 token 格式检查,并在 token 无效时打印带有推广性质的提示信息、网站和微信联系方式。这与声明的核心用途明显不一致。虽然 token 校验可能是某些技能的辅助实现,但当前代码片段完全没有体现任何视频下载、音视频解析、语音识别、字幕提取或总结处理行为,因此就该代码片段而言,实际行为与声明用途存在明显不匹配。

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · README.md (reported line 99)May include surrounding context.

md
4. 同时传入文件路径与任务ID,优先执行 `--id`,忽略 `--file`
5. 无自定义 prompt 时,默认完整转录视频全部文字

| 用户自然语言指令                                         | 自动生成命令                                                                                                |
| -------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 视频提取 https://example.com/video.mp4 中的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 把本地 /path/to/your/video.mp4 改成小红书风格的文案      | `node scripts/video2text/index.js --file "/path/to/your/video.mp4" --prompt "改写成小红书风格的文案"`       |

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill processes local files and public video URLs via a remote service, but the description does not present a clear upfront warning that user media is uploaded off-device. This is dangerous because users may unknowingly send sensitive meetings, interviews, or proprietary course content to a third party, creating privacy, confidentiality, and compliance exposure.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill claims it only communicates with guaikei.com, but examples and described behavior accept arbitrary external video URLs. This is dangerous because users and host agents may rely on the narrower network claim when deciding trust, while the actual workflow necessarily reaches untrusted third-party sources and can expose the environment to SSRF-like fetches, malicious media, or privacy leakage.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The documentation promises neutral errors without marketing or contact information, yet elsewhere embeds promotional contact details and directs users toward them. This is dangerous because it undermines trust boundaries and can normalize sending operational failures or sensitive context to out-of-band personal contacts, increasing phishing, social-engineering, and data-leak risk.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · SKILL.md (reported line 157)May include surrounding context.

md
## 7. 🗣️ 自然语言 → 命令(照这张表转)

| 用户说的话                                           | 就执行这条命令                                                                                              |
| ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 提取 https://example.com/video.mp4 里的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 总结这个视频的核心观点 https://example.com/video.mp4 | `node scripts/video2text/index.js --file "https://example.com/video.mp4" --prompt "总结这个视频的核心观点"` |

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This code sends the provided filename to an external API to obtain a presigned upload URL, and later functions send the uploaded video URL and prompt for remote processing. The file includes docstrings describing the API behavior, but there is no confirmation prompt, user-facing log, or warning to disclose that user content is being transmitted off-system.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The function sends a video URL to /api/video/id, initiating external analysis workflow, and getVideoText subsequently sends the analysis task ID and prompt to generate content. Although comments describe the API purpose, the code itself has no user disclosure, confirmation, or visible logging to warn that user media is being processed remotely.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This file contains user-facing natural-language strings in Chinese, beginning with the file header comment and continuing through thrown error messages and help output. Because the skill does not provide any language/locale choice or indicate that it is intentionally region-specific, it may violate language/locale policy by forcing a specific language on users.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

Error messages such as missing-value and unknown-option notices, as well as the generated help text, are emitted exclusively in Chinese. For a general-purpose utility file, this imposes a fixed locale on all users and there is no visible mechanism to select another language or confirm that Chinese-only output is acceptable.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

This downloader performs HTTP/HTTPS requests and writes the response to disk, and it can also delete partially downloaded files on stop or failure. While the class emits internal events, the file contains no confirmation prompt, print/log statement, or comment/docstring warning users about these safety-relevant behaviors.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This JavaScript file contains user-facing status/error messages and a doc comment entirely in Chinese, such as the download progress and failure notices. Because the file provides no language selection, fallback, or documented locale constraint, it appears to enforce a specific language on users without opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This code constructs and sends an outbound HTTPS POST request containing JSON data and a TOKEN header, which may transmit user or system data to an external service. There is no confirmation prompt or explicit user-facing disclosure in this file about the network transmission or inclusion of the token.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The file emits all user-facing status and warning text in Chinese, including the token misconfiguration warning and recovery instructions. There is no indication of language selection, opt-in, or a documented region-specific purpose, which creates a natural-language locale policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This code includes user-facing descriptions, help text, and runtime messages entirely in Chinese, which effectively forces a specific language for interaction. The file does not present any opt-in, alternative locale, or justification that the skill is intentionally limited to a Chinese-speaking context.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
73% confidence
Finding

The natural-language instructions, examples, trigger guidance, and operational text are entirely in Chinese, and the document does not offer an explicit language preference or opt-in for users who may expect another locale. Under the stated policy, forcing a specific language without user choice can be a locale-policy issue unless clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
99% confidence
Finding

The frontmatter defines the skill as guaikei-video-text-generator, but the convention section states the required canonical name, directory, package name, and H1 should be guaikei-video-to-doc. This is an internal documentation contradiction about the skill's identity and expected wiring.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

All human-readable comments and error messages in this file are written in Chinese, including operational labels passed into API requests. If these strings are surfaced to users, the skill effectively forces a single language without offering a locale choice or documenting a justified regional scope.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

The file uses hard-coded Chinese-language error strings for request failures and token problems, with no indication that the user can choose a language or that the skill is intentionally limited to a Chinese-speaking context. This creates a natural-language locale policy issue because it forces a specific language experience by default.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

The displayed banner text is entirely in Chinese, which suggests the skill presents its interface in a fixed language. Under the policy, forcing a specific language without offering a user choice or documenting a justified locale constraint is a natural-language policy issue.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.