Back to skill

Security audit

guaikei视频转文本

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly matches its video transcription purpose, but it uploads user media to cloud services and has identity and disclosure inconsistencies that users should review before installing.

Review this before installing if your videos may contain private, regulated, client, or business-sensitive content. Using it means local files or video links may be downloaded, uploaded to presigned cloud storage, and processed by the Guaikei service using your GUAIKEI_API_TOKEN. The publisher should correct the package name mismatch and make the actual upload destinations and error/contact behavior clearer.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (24)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

声明描述的是一个视频转写与文案生成技能,核心能力应包括接收视频链接/文件、提取音频或字幕、调用转写模型,并进一步生成总结或改写内容。但给出的代码片段只包含 args.js,一个通用命令行参数解析器:解析 --flag、别名、默认值、boolean、位置参数、帮助文本,以及参数校验和警告。这属于基础支撑工具代码,而非视频转写技能本身。就该代码片段可见行为而言,其主要用途与声明目的明显不一致,因此应判定为描述与代码行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的是“视频转文字 + 文案生成”技能,但提供的代码仅是下载工具模块。虽然下载远程视频链接可能是整个技能链路中的辅助步骤,但这段代码本身没有实现任何转写、字幕提取、语音识别、总结、改写或翻译功能,因此其实际行为与声明的核心能力明显不符。代码访问的资源主要是网络 URL 和本地文件系统,用于保存下载内容;这可以作为支持性实现存在,但如果用这段代码来代表该技能,则不能准确体现技能宣称的主要功能。因此应判定为描述与代码行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

声明描述的核心能力是“视频转文字/转写/总结生成”,而代码片段仅包含下载辅助函数和字节格式化、控制台日志功能。虽然下载视频文件或远程资源可能是该类技能的配套步骤,但从当前代码本身看,没有任何音视频解析、ASR 转写、字幕提取、文本清洗、总结生成或 Prompt 风格化处理能力。因此,这段代码的实际行为与声明的主要用途不一致,应判定为描述与行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Yes, this is a mismatch. The declared purpose is a video-to-text/transcription and content-generation skill, but the supplied code chunk does not handle video links, local video files, transcription, subtitle extraction, summarization, translation, or any cloud-model processing. Its behavior is limited to accessing the local filesystem to read package metadata and return the package name. That functionality is unrelated to the declared end-user purpose and indicates a materially different actual behavior for this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

声明描述的核心能力是视频转文字与文案生成,而该代码片段实际只是一个鉴权/配置辅助模块:检查 token 是否为指定格式,若无效则打印错误提示、官网地址和客服微信,并返回空字符串;若有效则返回 token。本片段没有任何与视频下载、音频提取、语音识别、字幕处理、总结改写或大模型生成相关的行为。虽然 token 校验可能是某类技能的配套实现细节,但就该片段本身而言,其实际行为与声明用途不一致,且包含未声明的营销信息输出,因此应判定为描述与行为不匹配。

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The document asserts strict identity consistency while presenting a canonical name that conflicts with the manifest name. Identity mismatches undermine provenance and increase the risk of users invoking or trusting the wrong package, which is dangerous in an agent-skill ecosystem where name-based routing and review matter.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · README.md (reported line 99)May include surrounding context.

md
4. 同时传入文件路径与任务ID,优先执行 `--id`,忽略 `--file`
5. 无自定义 prompt 时,默认完整转录视频全部文字

| 用户自然语言指令                                         | 自动生成命令                                                                                                |
| -------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 视频提取 https://example.com/video.mp4 中的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 把本地 /path/to/your/video.mp4 改成小红书风格的文案      | `node scripts/video2text/index.js --file "/path/to/your/video.mp4" --prompt "改写成小红书风格的文案"`       |

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill description does not give an upfront, prominent warning that user-provided local video/audio files and linked media are uploaded to a remote cloud service. This can cause users to disclose sensitive or regulated content without informed consent, which is particularly risky for meeting recordings, interviews, and private media.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The invocation guidance encourages broad automatic triggering on many video-related requests without first warning that content will be uploaded to a third-party cloud service. In an agent setting, that increases the chance of silent exfiltration of sensitive local files or private URLs under the guise of routine assistance.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill explicitly claims errors are neutral and free of marketing or contact links, yet the same document includes business-contact and promotional guidance. This inconsistency can facilitate social engineering by normalizing out-of-band contact during failure states, especially for users troubleshooting token or billing issues.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · SKILL.md (reported line 158)May include surrounding context.

md
## 7. 🗣️ 自然语言 → 命令(照这张表转)

| 用户说的话                                           | 就执行这条命令                                                                                              |
| ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 提取 https://example.com/video.mp4 里的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 总结这个视频的核心观点 https://example.com/video.mp4 | `node scripts/video2text/index.js --file "https://example.com/video.mp4" --prompt "总结这个视频的核心观点"` |

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This JavaScript file contains natural-language strings for errors, comments, and help output exclusively in Chinese, including user-facing CLI messages such as missing-argument errors and usage/help text. Per the policy, forcing a specific language without user opt-in is a locale/language policy violation unless the constraint is explicitly justified.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

This downloader performs outbound HTTP/HTTPS requests, writes downloaded content to disk, and can delete partially downloaded files on stop or failure. Although it emits internal events, the file contains no confirmation prompt, print/log disclosure, or explanatory comments/docstrings warning users about these safety-relevant actions.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest describes a skill focused on turning user-provided videos or links into transcripts and derived copy. This helper exposes a general-purpose URL download function with retry/resume/skip behavior for arbitrary files, which is broader than the stated transcription and copy-generation behavior and is not tied in code to video-only handling.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Several thrown error messages are written only in Chinese, including the credential-related guidance string. This imposes a fixed language on users with no opt-in or locale selection, which matches the language/locale policy violation category.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The code emits all user-facing status and promotional messages in Chinese only. This imposes a specific language on users without opt-in or an alternative locale, which is a natural-language policy violation under the language/locale rule.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The comment makes a concrete privacy and data-retention promise that is not enforced anywhere in this function, which only performs upload. In a skill that handles user videos, this can mislead operators and users into believing sensitive media is automatically deleted, increasing the risk of unintended retention, privacy violations, and compliance issues if downstream deletion is absent or fails.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The banner text is fixed in Chinese (视频文案智能提取助手), which indicates a language-specific user-facing experience without any visible opt-in or alternative locale handling in this file. The policy requires avoiding forced language/locale behavior unless the skill offers choice or clearly documents a justified regional constraint.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The README presents all user-facing instructions, warnings, and examples only in Chinese. Under the stated policy, forcing a specific language without user opt-in is a natural-language policy concern unless the locale restriction is explicitly justified, which is not documented here.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The conventions table says tags should be ≤30 and claims this skill has 30 个 at L225. Counting the actual tag entries in the frontmatter from L019-L048 yields more than 30 entries, so the documentation's claim about the current skill configuration is inconsistent with the manifest content.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

This manifest contains user-facing description, keywords, and author contact text in Chinese only, which effectively enforces a specific language for discovery and understanding. The file does not indicate that the skill is region-specific or provide any opt-in or alternative locale, so it may violate language/locale policy expectations.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The code automatically appends a skill identifier (skill_name) and sends the full JSON request body to a remote API, but this file contains no mechanism for user notice, consent, or minimization. In the context of a skill that processes user-supplied video links, local video files, and transcription content, this creates a real privacy and data-transmission risk because potentially sensitive media-derived content is sent off-box to a cloud service without any visible disclosure here.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The CLI descriptions, help text, examples, and user-facing messages are written in Chinese throughout the file, which effectively constrains interaction to a specific language. There is no indication that users can choose another language or opt in to this locale restriction.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.