Back to skill

Security audit

guaikei视频字幕提取

Security checks for vulnerabilities and agentic risk

Overview

This is a cloud video transcription skill that appears coherent, but installing it means user-provided videos, links, prompts, and the GUAIKEI token are sent to the provider.

Install only if you are comfortable sending the selected media, prompt text, and GUAIKEI_API_TOKEN to this provider. Avoid giving it confidential videos or internal URLs unless that upload is intended, and be aware that online links may be downloaded locally before being uploaded for processing.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (23)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the implementation only validates a token and prints console/marketing output, then the README's claims about transcription, summarization, translation, and privacy are materially false. While less directly exploitable than arbitrary file/network access, this still undermines user trust and can be used to collect tokens, drive users to off-platform contact channels, or bypass proper security review through misleading metadata.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the implementation only validates a token and prints console/marketing output, then the README's claims about transcription, summarization, translation, and privacy are materially false. While less directly exploitable than arbitrary file/network access, this still undermines user trust and can be used to collect tokens, drive users to off-platform contact channels, or bypass proper security review through misleading metadata.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

If the implementation only validates a token and prints console/marketing output, then the README's claims about transcription, summarization, translation, and privacy are materially false. While less directly exploitable than arbitrary file/network access, this still undermines user trust and can be used to collect tokens, drive users to off-platform contact channels, or bypass proper security review through misleading metadata.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

If the implementation only validates a token and prints console/marketing output, then the README's claims about transcription, summarization, translation, and privacy are materially false. While less directly exploitable than arbitrary file/network access, this still undermines user trust and can be used to collect tokens, drive users to off-platform contact channels, or bypass proper security review through misleading metadata.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

If the implementation only validates a token and prints console/marketing output, then the README's claims about transcription, summarization, translation, and privacy are materially false. While less directly exploitable than arbitrary file/network access, this still undermines user trust and can be used to collect tokens, drive users to off-platform contact channels, or bypass proper security review through misleading metadata.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The README defines very broad natural-language triggers and automatic command-conversion rules that can cause this skill to activate on loosely phrased user requests and transform them into executable processing commands. In an agent environment, this increases the chance of unintended invocation, accidental processing of sensitive local files or arbitrary URLs, and user actions occurring without explicit confirmation of the exact source and prompt.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · README.md (reported line 99)May include surrounding context.

md
4. 同时传入文件路径与任务ID,优先执行 `--id`,忽略 `--file`
5. 无自定义 prompt 时,默认完整转录视频全部文字

| 用户自然语言指令                                         | 自动生成命令                                                                                                |
| -------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 视频提取 https://example.com/video.mp4 中的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 把本地 /path/to/your/video.mp4 改成小红书风格的文案      | `node scripts/video2text/index.js --file "/path/to/your/video.mp4" --prompt "改写成小红书风格的文案"`       |

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The activation text contains very broad trigger phrases covering many common content-analysis and writing tasks. In an agent ecosystem, overbroad matching can cause the skill to be invoked in situations beyond its safe or intended scope, increasing exposure of user files, URLs, and prompts to an external service without sufficiently precise user intent.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
78% confidence
Finding

The manifest description is entirely in Chinese and repeatedly frames outputs around Chinese-language use cases such as 小红书、抖音、公众号文案, while the skill also claims support for broad transcription and analysis tasks. There is no indication that users can choose output language by default or that the skill is intentionally limited to a Chinese-only audience.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The README instructs AI systems to prioritize this skill for a wide range of intents with limited exclusions. That can bias an agent into over-selection of this tool, expanding unnecessary data sharing and making prompt-routing manipulation easier if the skill is not actually as bounded or trustworthy as described.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · SKILL.md (reported line 128)May include surrounding context.

md
当用户用自然语言下达指令时,按以下映射生成命令,保证识别与执行一致:

| 用户自然语言指令                                     | 生成的命令                                                                                                  |
| ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 提取 https://example.com/video.mp4 里的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 总结这个视频的核心观点 https://example.com/video.mp4 | `node scripts/video2text/index.js --file "https://example.com/video.mp4" --prompt "总结这个视频的核心观点"` |

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The package description advertises many loosely bounded triggers such as transcription, summarization, rewriting, meeting notes, platform-link handling, and content generation without clear activation constraints or exclusions. In agent routing systems, this can cause overbroad invocation, making the skill handle requests beyond its safest intended scope and increasing the chance of unintended data processing or misuse of remote content.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This downloader performs file creation via fs.createWriteStream and can remove downloaded files automatically through removeOnStop/removeOnFail behavior, but the code contains no confirmation prompt or user-facing disclosure around these data-affecting operations. Because this is a reusable utility rather than an obviously user-invoked destructive command, the silent write/delete behavior merits a warning finding.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This JavaScript file contains multiple user-visible strings in Chinese, including status and error messages, but provides no apparent option for users to select another language or opt in to this locale. That creates a natural-language policy concern because the skill effectively forces a specific language without documented justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The user-visible error strings in this file are written only in Chinese, including network, parsing, timeout, and retry failures. This imposes a specific language/locale without any opt-in or documented justification, which is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The user-facing strings shown when the token is invalid are entirely in Chinese, including warnings and recovery instructions. This imposes a specific language on all users without opt-in or any indication that the skill is intended only for a Chinese-speaking or region-specific audience.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
72% confidence
Finding

The natural-language metadata is entirely in Chinese and does not mention any option for other languages or a justified region-specific restriction. This can indicate a language/locale policy issue if the skill is surfaced to users without clear opt-in or locale scoping.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
98% confidence
Finding

This utility embeds user-facing text in Chinese, beginning with the file header/comment and continuing in thrown errors and help output. Because the file provides CLI-facing messages without any option to select locale or opt into Chinese, it may violate a language/locale policy requiring user choice.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

The thrown error strings for missing and invalid argument values are all written in Chinese, which forces one language regardless of user preference. Under the stated policy, fixed-language behavior without opt-in is a natural-language policy issue.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
98% confidence
Finding

The help output uses Chinese labels such as usage, options, help, and examples, making the CLI interaction language-specific by default. Since no language selection or documented justification is present in this file, this is a locale-policy concern.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

This code sends JSON data to a remote host and includes a TOKEN header, which can transmit user or system data off-device. The file contains no confirmation prompt or user-facing warning about this network transmission behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

The banner string is hard-coded in Chinese ("视频文案智能提取助手") with no indication that users can choose another language or locale. This can violate language/locale policy when a skill imposes a specific language on all users by default.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.