Back to skill

Security audit

guaikei视频内容转文稿提取

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent video transcription and content-generation skill that uploads user-selected media to a cloud API; users should treat submitted videos and prompts as third-party data sharing.

Install only if you are comfortable sending selected videos, prompts, task IDs, and the configured API token to GuaiKei's cloud service. Avoid using it for sensitive recordings unless you accept the provider's retention and privacy claims, and invoke it only for videos you intentionally want processed remotely.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (26)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

Token gating itself is normal, but the inclusion of promotional guidance and acquisition messaging inside an operational skill increases the chance of social-engineering style behavior and weakens the stated 'neutral errors/no marketing' trust boundary. In security-sensitive agent ecosystems, mixing execution instructions with vendor marketing is risky and misleading.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

Token gating itself is normal, but the inclusion of promotional guidance and acquisition messaging inside an operational skill increases the chance of social-engineering style behavior and weakens the stated 'neutral errors/no marketing' trust boundary. In security-sensitive agent ecosystems, mixing execution instructions with vendor marketing is risky and misleading.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

Token gating itself is normal, but the inclusion of promotional guidance and acquisition messaging inside an operational skill increases the chance of social-engineering style behavior and weakens the stated 'neutral errors/no marketing' trust boundary. In security-sensitive agent ecosystems, mixing execution instructions with vendor marketing is risky and misleading.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Token gating itself is normal, but the inclusion of promotional guidance and acquisition messaging inside an operational skill increases the chance of social-engineering style behavior and weakens the stated 'neutral errors/no marketing' trust boundary. In security-sensitive agent ecosystems, mixing execution instructions with vendor marketing is risky and misleading.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Token gating itself is normal, but the inclusion of promotional guidance and acquisition messaging inside an operational skill increases the chance of social-engineering style behavior and weakens the stated 'neutral errors/no marketing' trust boundary. In security-sensitive agent ecosystems, mixing execution instructions with vendor marketing is risky and misleading.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The README defines broad natural-language-to-command conversion rules that can cause the skill to activate on generic user requests such as summarization, extraction, analysis, or rewriting whenever a video URL or path is present. In an agent setting, this increases the chance of overbroad or unintended tool invocation, especially because the skill supports arbitrary --prompt content and remote URLs, which can send user-supplied media to an external service without sufficiently explicit user confirmation.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · README.md (reported line 99)May include surrounding context.

md
4. 同时传入文件路径与任务ID,优先执行 `--id`,忽略 `--file`
5. 无自定义 prompt 时,默认完整转录视频全部文字

| 用户自然语言指令                                         | 自动生成命令                                                                                                |
| -------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 视频提取 https://example.com/video.mp4 中的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 把本地 /path/to/your/video.mp4 改成小红书风格的文案      | `node scripts/video2text/index.js --file "/path/to/your/video.mp4" --prompt "改写成小红书风格的文案"`       |

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill requires access to an API token and network connectivity but does not declare an explicit tool scope such as permissions or allowed-tools. In an agent environment, missing scope declarations can cause overbroad execution privileges or make review and containment harder, especially for a skill that uploads user-provided media to a remote service.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill processes local files and public URLs by sending content to a remote service, but the user-facing description does not clearly warn that media may be uploaded off-device. This is a meaningful privacy and data-handling issue because users may submit sensitive recordings under the mistaken assumption of local-only processing.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The document claims errors will not include marketing content or extra contact/site references, but later sections include direct contact and promotional material. Security-wise, inconsistent trust statements are dangerous because they train users and agents to accept mixed operational and marketing output, which can mask phishing, support-scam, or data-exfiltration prompts.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · SKILL.md (reported line 157)May include surrounding context.

md
## 7. 🗣️ 自然语言 → 命令(照这张表转)

| 用户说的话                                           | 就执行这条命令                                                                                              |
| ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 提取 https://example.com/video.mp4 里的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 总结这个视频的核心观点 https://example.com/video.mp4 | `node scripts/video2text/index.js --file "https://example.com/video.mp4" --prompt "总结这个视频的核心观点"` |

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The package description is written entirely in Chinese and presents the skill's behavior only in that language. For an all-file-types policy check, this can constitute a locale/language policy violation because it does not offer a language choice or document a justified region-specific constraint.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This code sends user-provided file path information to the /api/video/presign endpoint and later transmits the uploaded video URL to /api/video/id. Although the functions have developer-facing comments, there is no confirmation prompt, print/log disclosure, or other user-facing warning in the file about transmitting user data to a remote API.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This JavaScript file contains natural-language comments and user-visible status/error messages exclusively in Chinese, including cleanup notifications and disk-space warnings. Because the file does not offer a language choice or document that the skill is intentionally region-specific, it appears to force a specific language/locale, which matches the policy-violation category.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This JavaScript file contains natural-language comments and runtime status messages exclusively in Chinese, including the function docstring and all download/retry/error messages. Because the skill does not provide any opt-in, fallback, or documented justification for a Chinese-only locale, it violates the language-choice policy for natural-language content.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The file emits all user-facing status and promotional text in Chinese, including warnings and operational messages, with no indication that the skill is China-specific or that users can select another language. This creates a natural-language locale policy issue because users are forced into a specific language without opt-in.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The comment makes strong privacy and data-handling guarantees such as non-retention, non-sharing, and automatic deletion, but the function only performs an HTTPS upload and contains no technical controls to enforce deletion, retention limits, destination validation beyond HTTPS, or restricted downstream use. In a video-transcription skill that handles potentially sensitive user media, this creates a trust gap where users and integrators may believe protections exist when they are not actually implemented.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

The user-facing banner string is entirely in Chinese and presents the tool as a Chinese-language assistant, with no indication elsewhere in this file that users can select another language or opt in to this locale. That can conflict with language/locale policy requirements that prohibit forcing a specific language without user choice.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The user-facing flag descriptions, help text, and examples are all written in Chinese, which effectively forces a specific language for interaction. The file does not offer an opt-in language choice or explain that the skill is intentionally limited to a Chinese-speaking audience or region-specific use case.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The markdown includes preset output modes such as '小红书文案', '抖音短标题', and '公众号文章' that implicitly force Chinese-language, China-platform-specific output styles. While the file later mentions language adaptation, these templates and examples do not consistently state that language/style should follow user preference, which can conflict with locale-choice expectations.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
98% confidence
Finding

Conflicting skill names and paths undermine the document's own integrity guarantees and can cause an agent or operator to invoke the wrong package or trust the wrong artifact. While not directly exploitable alone, this weakens provenance, reviewability, and safe deployment controls.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The file's comments and error/status strings are written exclusively in Chinese, which indicates a fixed language choice without offering user selection or documenting a justified locale constraint. This can violate language/locale policy when the skill is expected to operate for general users without explicit opt-in to that language.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

This file contains user-facing natural-language strings in Chinese, beginning with the top-level description and continuing throughout errors and help text. Because the skill does not offer localization, user choice, or a documented justification for a Chinese-only interface, it may violate language/locale policy requirements.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

Natural-language strings shown to users, including validation errors and help text, are all emitted in Chinese across the parser and help generator. The file does not provide user opt-in, fallback language behavior, or any documented regional justification for enforcing this locale.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The string forces a Chinese-language error and support instruction, including contact guidance, with no indication of locale selection or user opt-in. This can violate language/locale policy when the skill is not clearly documented as Chinese-only or region-specific.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.