Back to skill

Security audit

guaikei语音转文字

Security checks for vulnerabilities and agentic risk

Overview

This skill is a cloud video/audio transcription tool that uploads user-specified media to a remote service, and the inspected behavior is mostly disclosed and aligned with that purpose.

Install only if you are comfortable sending the selected video/audio files, media URLs, prompts, filenames, and API token-authenticated requests to the GuaiKei cloud service and its HTTPS upload storage. Avoid using it for confidential meetings, private recordings, or copyrighted media unless the provider's retention and privacy claims meet your needs.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (16)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill claims constrained behavior but apparently supports general network download flows, local file write/delete, resume, redirects, and retries. Those capabilities materially expand the attack surface: user-supplied URLs can lead to unreviewed remote fetching and local filesystem modification beyond a narrowly described transcription action.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill claims constrained behavior but apparently supports general network download flows, local file write/delete, resume, redirects, and retries. Those capabilities materially expand the attack surface: user-supplied URLs can lead to unreviewed remote fetching and local filesystem modification beyond a narrowly described transcription action.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill claims constrained behavior but apparently supports general network download flows, local file write/delete, resume, redirects, and retries. Those capabilities materially expand the attack surface: user-supplied URLs can lead to unreviewed remote fetching and local filesystem modification beyond a narrowly described transcription action.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill claims constrained behavior but apparently supports general network download flows, local file write/delete, resume, redirects, and retries. Those capabilities materially expand the attack surface: user-supplied URLs can lead to unreviewed remote fetching and local filesystem modification beyond a narrowly described transcription action.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Although the README mentions cloud processing and privacy claims, it does not prominently warn users at the point of use that submitted local videos and third-party links are sent to a remote service controlled outside the local environment. Users may mistakenly believe processing is local or equivalent to offline tools, leading to inadvertent disclosure of confidential recordings, meetings, or copyrighted content.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The README defines very broad natural-language trigger phrases and auto-conversion rules that can map ordinary user requests directly into a command that uploads local files or remote video URLs for cloud processing. This creates a prompt/skill overreach risk: users may invoke the skill unintentionally, causing unintended data transfer or processing of sensitive media.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · README.md (reported line 99)May include surrounding context.

md
4. 同时传入文件路径与任务ID,优先执行 `--id`,忽略 `--file`
5. 无自定义 prompt 时,默认完整转录视频全部文字

| 用户自然语言指令                                         | 自动生成命令                                                                                                |
| -------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 视频提取 https://example.com/video.mp4 中的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 把本地 /path/to/your/video.mp4 改成小红书风格的文案      | `node scripts/video2text/index.js --file "/path/to/your/video.mp4" --prompt "改写成小红书风格的文案"`       |

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill requires access to a sensitive environment variable (GUAIKEI_API_TOKEN) but does not declare any explicit tool scope or permissions boundary. That weakens reviewability and policy enforcement, because the runtime can access secrets without the manifest clearly constraining or documenting that capability.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The trigger description is extremely broad and covers many generic content-processing tasks, increasing the chance the skill activates for requests beyond simple speech-to-text. In an agent ecosystem, overbroad activation can cause unintended exfiltration of local files or remote media to a third-party service when a more limited or local tool would have been appropriate.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The guidance says to prioritize this skill when any matching intent appears, but the non-applicable cases are not strong enough to prevent over-selection. That can route broad analysis or writing tasks into a networked media-upload workflow, exposing data to a third-party service without sufficiently precise user intent.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · SKILL.md (reported line 128)May include surrounding context.

md
当用户用自然语言下达指令时,按以下映射生成命令,保证识别与执行一致:

| 用户自然语言指令                                     | 生成的命令                                                                                                  |
| ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 提取 https://example.com/video.mp4 里的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 总结这个视频的核心观点 https://example.com/video.mp4 | `node scripts/video2text/index.js --file "https://example.com/video.mp4" --prompt "总结这个视频的核心观点"` |

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This code file contains multiple natural-language strings for errors, warnings, and help output entirely in Chinese, such as the module description, thrown errors, and generated CLI help text. Because the skill forces a specific language for user interaction without any opt-in or documented locale limitation, it violates the language/locale policy criterion.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

This code builds and sends an HTTPS POST request containing arbitrary JSON payload data plus a TOKEN header. There is input validation and retry logging, but no confirmation prompt, disclosure message, or warning that user/system data and credentials will be transmitted to a remote host.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

This code emits all user-facing status and warning text in Chinese only, including configuration guidance and recovery instructions. The file provides no indication of user opt-in, locale detection, or an alternative language path, which is a natural-language locale policy violation under the stated rules.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

Multiple thrown error messages are hard-coded in Chinese, and the file provides no opt-in, locale selection, or justification for restricting output language. This can violate language/locale policy when the skill is expected to operate for users in other languages.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

The banner text is hard-coded in Chinese, which imposes a specific language on all users regardless of their locale or preferences. This is a natural-language policy concern because the file provides no mechanism to choose or opt into the language.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.