Back to skill

Security audit

guaikei视频秒出字幕文字

Security checks for vulnerabilities and agentic risk

Overview

This appears to be a legitimate cloud video-to-text tool, but using it sends the chosen video or video link to GuaiKei's service.

Install only if you are comfortable giving GuaiKei a token and sending selected videos, public video links, prompts, filenames, and task IDs to its cloud service. Do not use it for confidential, regulated, internal-only, or access-controlled media unless your organization has approved that provider and its retention claims.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (25)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

A skill that claims to transcribe videos but appears to implement token validation and promotional/configuration flows instead is misrepresenting its core behavior. Even if not directly exploitable by code execution, this erodes trust and can socially engineer users into disclosing credentials or accepting unnecessary external service use.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

A skill that claims to transcribe videos but appears to implement token validation and promotional/configuration flows instead is misrepresenting its core behavior. Even if not directly exploitable by code execution, this erodes trust and can socially engineer users into disclosing credentials or accepting unnecessary external service use.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

A skill that claims to transcribe videos but appears to implement token validation and promotional/configuration flows instead is misrepresenting its core behavior. Even if not directly exploitable by code execution, this erodes trust and can socially engineer users into disclosing credentials or accepting unnecessary external service use.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

A skill that claims to transcribe videos but appears to implement token validation and promotional/configuration flows instead is misrepresenting its core behavior. Even if not directly exploitable by code execution, this erodes trust and can socially engineer users into disclosing credentials or accepting unnecessary external service use.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The documentation encourages direct processing of external video links and cloud-based transcription but does not present a clear, prominent warning about the security and privacy risks of transmitting potentially sensitive media to a third-party service. Even though it claims encryption and deletion, users are not adequately informed about trust boundaries, remote fetching risks, or the possibility that private/internal URLs or confidential local videos could be sent off-device.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The README defines very broad natural-language trigger and auto-conversion rules that can cause the skill to activate on loosely matched requests and automatically transform user input into executable command arguments such as --file, --id, and --prompt. In a skill that uploads local files or fetches external URLs for cloud processing, overbroad invocation increases the chance of unintended data submission, accidental reuse of prior task IDs, or processing of sensitive media without sufficiently explicit user confirmation.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · README.md (reported line 99)May include surrounding context.

md
4. 同时传入文件路径与任务ID,优先执行 `--id`,忽略 `--file`
5. 无自定义 prompt 时,默认完整转录视频全部文字

| 用户自然语言指令                                         | 自动生成命令                                                                                                |
| -------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 视频提取 https://example.com/video.mp4 中的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 把本地 /path/to/your/video.mp4 改成小红书风格的文案      | `node scripts/video2text/index.js --file "/path/to/your/video.mp4" --prompt "改写成小红书风格的文案"`       |

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill declares access to an environment secret (GUAIKEI_API_TOKEN) and describes networked/cloud processing, but it does not define an explicit tool scope such as permissions or allowed-tools. In an agent environment, missing scope boundaries weakens least-privilege controls and can let the runtime grant broader capabilities than users or reviewers expect.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill processes local files and remote video links in a cloud service, but the warning about user content being uploaded is not sufficiently clear and prominent. This is dangerous because users may unintentionally send sensitive local media or private links off-device without informed consent.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The document says error messages must be neutral and free of marketing or website links, but elsewhere provides contact and commercial guidance tied to error handling. This inconsistency is dangerous because it opens the door to phishing-like or coercive recovery flows, especially when users are troubleshooting authentication or upload failures.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · SKILL.md (reported line 157)May include surrounding context.

md
## 7. 🗣️ 自然语言 → 命令(照这张表转)

| 用户说的话                                           | 就执行这条命令                                                                                              |
| ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 提取 https://example.com/video.mp4 里的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 总结这个视频的核心观点 https://example.com/video.mp4 | `node scripts/video2text/index.js --file "https://example.com/video.mp4" --prompt "总结这个视频的核心观点"` |

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The functions submit a video URL and later a free-form prompt to external API endpoints, which may expose user content or metadata to a remote service. The file contains docstrings for developers, but no confirmation, print/log disclosure, or explicit user warning about this data transmission.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This code initiates HTTP/HTTPS requests, writes downloaded content to disk, and may delete files on stop or failure, but it contains no confirmation prompt, print/log statement, or explanatory comment/docstring warning users about these effects. For a reusable utility, these safety-relevant actions are not disclosed within the file itself.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The string literal forces a Chinese-language error message and directs users to contact via WeChat/phone, with no indication that language or locale is configurable. This can violate language/locale policy because the skill imposes a specific language and region-specific support channel without documented user choice or justification in this file.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The file emits all user-facing warning and status messages exclusively in Chinese, including the invalid-token warning and follow-up guidance. This imposes a specific language on all users without offering a language selection or documenting a justified locale restriction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This code presents all user-facing descriptions, examples, and runtime guidance in Chinese, including flag descriptions and help text. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation unless the locale restriction is clearly justified, which is not documented here.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The manifest description is entirely in Chinese and presents the skill's invocation phrases only in Chinese, which can impose a locale expectation without explicit user opt-in. The policy allows locale constraints when clearly documented and justified, but this file does not clearly state that the skill is intended only for Chinese-speaking users or provide a language choice.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

L141、L348、L445 都说明 --id last 只保留 24 小时,但 L470 又写“中途退出也不怕,一小时内可用任务ID(或 24 小时内用 --id last)回来取结果”。这会误导实现或调用方对复用窗口作出相互矛盾的判断,属于文档意图与行为契约不一致。

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

This is a manifest file, so vague-trigger review applies. The description lists many broad use cases like turning videos into summaries, scripts, and social-media copy, but does not define any explicit invocation phrases, scope limits, or exclusion conditions, which can increase the chance of unintended matching in systems that rely on manifest text for routing.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

All visible error messages and operation labels in this file are written in Chinese, with no indication that users can choose another language or locale. This can violate language/locale policy when a skill forces a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

This code transmits the provided filename to an API endpoint to obtain a presigned upload URL, but there is no confirmation prompt, log, or other user-visible disclosure in the file. Because file paths can reveal user or system information, this network transmission should be disclosed unless warning is provided elsewhere in markdown documentation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

This JavaScript file contains user-facing natural-language strings such as error messages and help output entirely in Chinese. Under the policy, forcing a specific language without user opt-in or clear justification is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

This file contains natural-language text only in Chinese in both the function comment and user-visible log/error messages. Because the skill does not offer language selection or indicate that it is intentionally limited to Chinese users, it can violate a language/locale policy requiring user choice or explicit justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

This JavaScript file contains multiple user-visible strings such as download completion, retry, resume, and failure messages written only in Chinese. That can violate a language/locale policy when users are not given any opt-in or localization choice.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The user-facing banner string is entirely in Chinese, which indicates a language-specific skill presentation. In this file there is no visible option, prompt, or documentation allowing the user to choose another language, which can violate a language/locale policy requiring opt-in or choice.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.