Back to skill

Security audit

guaikei-video-copy

Security checks for vulnerabilities and agentic risk

Overview

The skill is generally aligned with video transcription, but it can send user videos to cloud storage and has inconsistent disclosure around network destinations, retention, and token-error messaging.

Install only if you are comfortable sending the selected video or downloaded public video to GuaiKei/cloud storage for processing. Do not use it for sensitive meetings, regulated data, confidential interviews, or private media unless you have approval and trust the provider's retention claims. Expect it to store a last-task ID and temporary downloads locally for reuse/cleanup, and be aware that token errors may show off-platform contact and website information.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (32)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill asserts that error messages should be neutral and not include marketing links, yet the finding indicates token-gating behavior that outputs website/contact information. Embedding marketing or off-platform contact prompts in failure paths can manipulate users, obscure operational errors, and expand trust to external channels not required for the task.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill asserts that error messages should be neutral and not include marketing links, yet the finding indicates token-gating behavior that outputs website/contact information. Embedding marketing or off-platform contact prompts in failure paths can manipulate users, obscure operational errors, and expand trust to external channels not required for the task.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill asserts that error messages should be neutral and not include marketing links, yet the finding indicates token-gating behavior that outputs website/contact information. Embedding marketing or off-platform contact prompts in failure paths can manipulate users, obscure operational errors, and expand trust to external channels not required for the task.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill asserts that error messages should be neutral and not include marketing links, yet the finding indicates token-gating behavior that outputs website/contact information. Embedding marketing or off-platform contact prompts in failure paths can manipulate users, obscure operational errors, and expand trust to external channels not required for the task.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill asserts that error messages should be neutral and not include marketing links, yet the finding indicates token-gating behavior that outputs website/contact information. Embedding marketing or off-platform contact prompts in failure paths can manipulate users, obscure operational errors, and expand trust to external channels not required for the task.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The README promotes cloud-based processing and direct handling of public video links and local files, but does not prominently warn users that content is transmitted to a remote third-party service for analysis. This can lead users or higher-level agents to send sensitive videos, meeting recordings, or internal URLs off-device without informed consent, creating confidentiality and compliance risk.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The README defines broad natural-language-to-command conversion rules that can cause an agent to treat loosely phrased user text as executable parameters, including arbitrary URLs, local file paths, and free-form prompts. In an agent setting, weak trigger boundaries increase the chance of unintended remote fetches, accidental disclosure of local files, or prompt injection into downstream processing.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · README.md (reported line 99)May include surrounding context.

md
4. 同时传入文件路径与任务ID,优先执行 `--id`,忽略 `--file`
5. 无自定义 prompt 时,默认完整转录视频全部文字

| 用户自然语言指令                                         | 自动生成命令                                                                                                |
| -------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 视频提取 https://example.com/video.mp4 中的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 把本地 /path/to/your/video.mp4 改成小红书风格的文案      | `node scripts/video2text/index.js --file "/path/to/your/video.mp4" --prompt "改写成小红书风格的文案"`       |

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This skill processes local video/audio and says it uses a cloud model, but the warning about user content being uploaded to a remote service is not sufficiently explicit at the point of use. Users may unknowingly transmit sensitive meetings, interviews, or proprietary media off-device, creating privacy and compliance risk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The documentation mandates a Chinese H1 format ('中文名 +(name)') and elsewhere standardizes terminology and examples entirely around Chinese usage, without presenting user choice or documenting a justified locale restriction. Under the language/locale policy rule, hard-coding a specific language convention without opt-in can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill handles local video files and public video URLs by uploading content to a cloud service, but the user-facing description does not make that data transfer explicit at the point of use. This can cause users to unknowingly send sensitive recordings, meeting content, or personal data off-device, creating a real privacy and data-governance risk.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Multiple sections say --id last and local task reuse are valid for 24 hours, but the FAQ states users can come back within one hour using the task ID. This is an active contradiction in documented behavior that can mislead agents about when reuse is supported.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · SKILL.md (reported line 158)May include surrounding context.

md
## 7. 🗣️ 自然语言 → 命令(照这张表转)

| 用户说的话                                           | 就执行这条命令                                                                                              |
| ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 提取 https://example.com/video.mp4 里的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 总结这个视频的核心观点 https://example.com/video.mp4 | `node scripts/video2text/index.js --file "https://example.com/video.mp4" --prompt "总结这个视频的核心观点"` |

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · guaikei-video2text-1.0.1/SKILL.md (reported line 158)May include surrounding context.

md
## 7. 🗣️ 自然语言 → 命令(照这张表转)

| 用户说的话                                           | 就执行这条命令                                                                                              |
| ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 提取 https://example.com/video.mp4 里的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 总结这个视频的核心观点 https://example.com/video.mp4 | `node scripts/video2text/index.js --file "https://example.com/video.mp4" --prompt "总结这个视频的核心观点"` |

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest description is written entirely in Chinese and explicitly promises Chinese/English translation, but it does not indicate that users may choose their preferred interaction language or locale. This can violate language/locale policy when a skill appears to assume a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This JavaScript file uses Chinese exclusively in comments, thrown error messages, and generated help output, which imposes a specific language on users of the CLI. Under the policy, locale or language constraints should either be optional, user-selectable, or clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The downloader deletes files via fs.unlink when stopping and also removes files on failure, yet this code does not include any user-facing warning, confirmation, or comment explaining that partially downloaded files may be deleted automatically. Automatic deletion is a potentially destructive action and should be disclosed clearly.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

This code performs network requests to arbitrary URLs and writes the retrieved content to disk, but the file contains no confirmation prompt, print/log statement, or explanatory comment warning users that remote data will be fetched and saved locally. Because the operation affects both network exposure and the local filesystem, some explicit disclosure is expected unless it is clearly documented elsewhere, which is not visible in this file.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The code emits user-facing error messages exclusively in Chinese, including parsing, timeout, network, and token-related failures. This imposes a specific language on all users without offering any locale choice or documented justification, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill emits all user-facing warning and status messages only in Chinese, including the invalid-token warning and promotional guidance. This imposes a specific language on users without opt-in or an alternative locale, which matches the language/locale policy violation category.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The docstring promises that uploaded videos will not be retained and will be automatically deleted after transcription, but this function only performs an upload and contains no deletion, lifecycle enforcement, or verification of remote retention policy. That gap can mislead users into sharing sensitive recordings under false assumptions, increasing privacy, compliance, and trust risks.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The function streams a local file to a remote presigned OSS URL, which is an external data transfer of user-provided video content. The skill description emphasizes transcription outcomes but does not clearly disclose that local files are uploaded to remote storage, creating a meaningful privacy and data-handling risk, especially because videos may contain sensitive audio, faces, meetings, or proprietary material.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

User-facing descriptions, help text, and error/output strings in this file are written exclusively in Chinese, which effectively forces a specific language for interaction. The file does not provide an opt-in language selection or justify a locale-specific restriction, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The script uploads user-supplied local or downloaded video files to a remote service and only prints a generic status message about uploading, without an explicit consent prompt or clear warning that potentially sensitive media will leave the local machine. Because this skill is specifically designed to process videos that may contain private conversations, meetings, classes, or interviews, users could unintentionally transmit sensitive content to third-party infrastructure.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.