Tp4
- Category
- MCP Tool Poisoning
- Confidence
- 99% confidence
- Finding
声明描述的是一个面向视频内容处理的技能,核心能力应包括接收视频/链接、解析媒体内容、转写、提炼和生成文案。但给出的代码片段仅实现了通用参数解析与帮助信息构建,属于基础 CLI 支撑模块。按照评估标准,支持性实现细节本身不应被误判;但这里的问题在于,当前代码片段完全没有体现声明中的任何核心业务行为,因此该代码与声明用途存在明显不一致。
- Content
Security audit
Security checks for vulnerabilities and agentic risk
This skill appears to do the advertised video transcription work, but its network and data-flow scope is broader than its documentation claims.
Install only if you are comfortable sending selected videos, prompts, task IDs, and related metadata to the Guaikei cloud service. Avoid using it on confidential recordings unless you trust the provider's retention promises, and do not feed it untrusted or internal-network URLs because the downloader may fetch and upload whatever that URL returns.
声明描述的是一个面向视频内容处理的技能,核心能力应包括接收视频/链接、解析媒体内容、转写、提炼和生成文案。但给出的代码片段仅实现了通用参数解析与帮助信息构建,属于基础 CLI 支撑模块。按照评估标准,支持性实现细节本身不应被误判;但这里的问题在于,当前代码片段完全没有体现声明中的任何核心业务行为,因此该代码与声明用途存在明显不一致。
该代码片段的核心功能是下载远程资源并管理下载状态,是一个底层下载工具类。虽然视频链接解析场景可能需要下载视频作为前置步骤,但当前代码没有任何音视频解析、语音识别、字幕提取、内容总结、文案加工、会议纪要生成或自定义 Prompt 处理逻辑。其实际主功能与声明的“视频转写与内容加工”存在明显偏差,因此应判定为描述与行为不匹配。
The declared description describes a full video-to-text and content-processing capability. The supplied code chunk does not implement any media ingestion, transcription, summarization, link handling, cloud parsing, or prompt-based transformation. It only accesses a local metadata file (package.json) to obtain the skill/package name. This is materially unrelated to the declared primary purpose, so it is a clear mismatch.
The declared purpose is a video-to-text/content-processing skill, but the supplied code chunk is solely an authentication/configuration helper for validating an API token. While token handling could be a supporting utility in a larger skill, this specific chunk does not exhibit the declared transcription, subtitle extraction, summarization, translation, or link/file processing behavior. It also includes undeclared promotional messaging related to purchasing or obtaining a token. Therefore, this code chunk does not accurately represent the declared functional purpose.
Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.
4. 同时传入文件路径与任务ID,优先执行 `--id`,忽略 `--file`
5. 无自定义 prompt 时,默认完整转录视频全部文字
| 用户自然语言指令 | 自动生成命令 |
| -------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 视频提取 https://example.com/video.mp4 中的文字 | `node scripts/video2text/index.js --file "https://example.com/video.mp4"` |
| 把本地 /path/to/your/video.mp4 改成小红书风格的文案 | `node scripts/video2text/index.js --file "/path/to/your/video.mp4" --prompt "改写成小红书风格的文案"` |
Without declared permissions the skill's intent is opaque and cannot be validated.
The skill processes local files and remote URLs through a cloud service, but the description does not prominently warn users that their media content is uploaded off-device for remote processing. This creates a real privacy and consent risk, especially for meetings, interviews, courses, or other sensitive recordings that may contain confidential or personal data.
The documentation claims the skill only communicates with guaikei.com, but the examples and described behavior explicitly accept arbitrary external video URLs. This can mislead users and orchestrators about the real data flow, increasing the risk of unintended requests to attacker-controlled hosts, privacy leakage, or SSRF-like fetch behavior if URL handling is not tightly constrained.
Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.
## 7. 🗣️ 自然语言 → 命令(照这张表转)
| 用户说的话 | 就执行这条命令 |
| ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 提取 https://example.com/video.mp4 里的文字 | `node scripts/video2text/index.js --file "https://example.com/video.mp4"` |
| 总结这个视频的核心观点 https://example.com/video.mp4 | `node scripts/video2text/index.js --file "https://example.com/video.mp4" --prompt "总结这个视频的核心观点"` |
The I/O contract says 'stdout 只有最终文案' and later reiterates that stdout contains only the final copy with no extra data. But sections 11.3 and 16.3 instruct users to '记下首次输出的任务ID' and return later with that ID, which implies the task ID is emitted somewhere as part of output; this contradicts the stated output contract unless the ID is persisted only locally and never shown.
The natural-language content in comments and error messages throughout the file is exclusively Chinese, such as the docstrings and thrown error text. This can violate a language/locale policy when the skill does not offer a language choice or document that it is intentionally region-specific.
The downloader performs network requests to arbitrary attacker-controlled URLs with support for redirects, which can enable server-side request forgery if upstream code passes untrusted URLs into this utility. In this skill context, accepting platform links and remote video sources makes the risk more relevant because an attacker could supply internal or sensitive endpoints instead of expected media hosts.
This downloader creates and writes to files on disk, and elsewhere in the file it can remove those files on stop or failure. Although the class emits internal events, there is no confirmation prompt, print/log disclosure, or explanatory comment/docstring warning users that invoking it will modify and potentially delete local files.
This JavaScript file contains natural-language content in comments and runtime log messages exclusively in Chinese, such as the docstring and multiple status/error strings. Because the skill does not offer a language choice or document a justified locale restriction, it risks violating language/locale policy for users expecting another language.
The 401/403 error path returns a Chinese-only message instructing the user to check an environment variable or contact a phone/WeChat number. This imposes a specific language on users without any opt-in or fallback, which matches the locale/language policy concern.
The user-facing messages printed on invalid token are entirely in Chinese, which imposes a specific language on users without opt-in or fallback. This is a natural-language policy concern because the file provides no indication that the skill is intentionally region-specific or that alternative locales are supported.
The code and its public-facing comment make a privacy and data-lifecycle promise that uploaded videos are only used for transcription and automatically deleted, but this module implements only upload behavior and contains no technical enforcement of deletion, retention limits, or use restrictions. In a skill handling potentially sensitive video content, this creates a trust gap where user data may persist in object storage or be accessible beyond the stated purpose, leading to privacy, compliance, and contractual risk.
The user-facing flag descriptions, help text, examples, and runtime messages are written in Chinese only. This imposes a specific language on all users without any opt-in or indication that the skill is intentionally region-specific, which matches the language/locale policy violation criteria.
The frontmatter declares the skill name as 'guaikei-video-speech-recognition', but the specification section claims the canonical name, directory, package, and H1 should be 'guaikei-video-to-doc'. This is an active contradiction in documentation about the skill's identity, not merely an omission.
The natural-language support instructions direct users to Chinese-language contact channels such as WeChat and a Chinese support workflow, but do not indicate that the skill is region-specific or offer alternatives for other locales. This can violate language/locale policy where user-facing operational guidance is forced into one locale without opt-in.
The package description lists many general-purpose actions such as transcription, summarization, copywriting, and script generation, but does not define when the skill should activate or what exact requests should map to it. In a manifest-scoped description, this breadth can make routing ambiguous and increase the chance of unintended invocation for common content-creation requests.
The description, keywords, and author metadata strongly imply a Chinese-language-only experience, but the manifest does not state that the skill is intended only for Chinese-speaking users or offer a language choice. That can create a locale policy issue if the surrounding platform expects skills not to force a language without explicit opt-in or justification.
This code sends potentially sensitive inputs to remote API endpoints, including an API token and a filename/path at L21-L27, and later a video URL and prompt in similar calls. While the functions have internal docstrings, there is no user-facing confirmation, warning, or visible disclosure in this file that user data will be transmitted over the network.
This file embeds user-facing natural-language strings in Chinese, beginning with the top-level description and continuing throughout error and help output. Because the skill forces a specific language for user interaction without opt-in or justification, it creates a locale policy concern under the natural-language policy rules.
This JavaScript file contains its documentation comment and all user-visible log/error strings exclusively in Chinese. Under the policy rule for natural-language violations, this is a locale/language constraint without any visible opt-in, fallback, or justification that the skill is region-specific.
No suspicious patterns detected.