Back to skill

Security audit

guaikei口播视频转稿

Security checks for vulnerabilities and agentic risk

Overview

This skill is a disclosed cloud video transcription tool, but users should understand that local videos, video URLs, prompts, task IDs, and the API token are sent to the provider's service.

Install only if you are comfortable sending video content, prompts, task IDs, and your GUAIKEI_API_TOKEN to the guaikei cloud workflow. Confirm before uploading local or sensitive media, avoid private/internal URLs, and do not rely solely on the package's retention and deletion claims unless the provider's terms meet your requirements.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (28)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the implementation does not actually perform the promised transcription/link processing and instead accesses local filesystem/package metadata, the skill is misrepresenting what data it touches. That is dangerous because users may provide sensitive local paths or expect cloud transcription while the skill performs unrelated local inspection, creating a trust and privacy boundary violation.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the implementation does not actually perform the promised transcription/link processing and instead accesses local filesystem/package metadata, the skill is misrepresenting what data it touches. That is dangerous because users may provide sensitive local paths or expect cloud transcription while the skill performs unrelated local inspection, creating a trust and privacy boundary violation.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the implementation does not actually perform the promised transcription/link processing and instead accesses local filesystem/package metadata, the skill is misrepresenting what data it touches. That is dangerous because users may provide sensitive local paths or expect cloud transcription while the skill performs unrelated local inspection, creating a trust and privacy boundary violation.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

If the implementation does not actually perform the promised transcription/link processing and instead accesses local filesystem/package metadata, the skill is misrepresenting what data it touches. That is dangerous because users may provide sensitive local paths or expect cloud transcription while the skill performs unrelated local inspection, creating a trust and privacy boundary violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documentation emphasizes convenience and cloud processing but does not place an explicit, front-and-center warning near usage examples that video files and public links are transmitted to a remote third-party service. Users may reasonably assume local-only handling, especially when passing local file paths, leading to uninformed disclosure of potentially sensitive audio/video content.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The README teaches broad natural-language trigger patterns such as '总结一下这个视频' or '用刚才分析的视频提取所有金句' that can overlap with ordinary conversational input. In an agent setting, this increases the chance of unintended tool execution, causing accidental upload of local files or remote URLs to the service and unintended use of prior task IDs.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · README.md (reported line 99)May include surrounding context.

md
4. 同时传入文件路径与任务ID,优先执行 `--id`,忽略 `--file`
5. 无自定义 prompt 时,默认完整转录视频全部文字

| 用户自然语言指令                                         | 自动生成命令                                                                                                |
| -------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 视频提取 https://example.com/video.mp4 中的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 把本地 /path/to/your/video.mp4 改成小红书风格的文案      | `node scripts/video2text/index.js --file "/path/to/your/video.mp4" --prompt "改写成小红书风格的文案"`       |

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill encourages sending user-provided local files or URLs to a cloud service, but the invocation guidance does not place a prominent privacy warning at the decision point. Users may unknowingly upload sensitive local media or internal links, causing data disclosure to a third party outside the local environment.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The document states that error messages should be neutral and must not include marketing copy, contact information, or website links. However, later sections explicitly direct users to the official website for token setup and include business/contact details, which undermines that stated intent and creates contradictory operator guidance around user-facing failures.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · SKILL.md (reported line 157)May include surrounding context.

md
## 7. 🗣️ 自然语言 → 命令(照这张表转)

| 用户说的话                                           | 就执行这条命令                                                                                              |
| ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 提取 https://example.com/video.mp4 里的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 总结这个视频的核心观点 https://example.com/video.mp4 | `node scripts/video2text/index.js --file "https://example.com/video.mp4" --prompt "总结这个视频的核心观点"` |

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The package description states the skill can convert video from a local file or public links such as Douyin/Xiaohongshu, but it does not define specific invocation phrases, boundaries, or exclusion conditions. In a manifest file, this broad natural-language scope can overlap with many common user requests about videos and make activation conditions unclear.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The default options enable removeOnStop and removeOnFail, and later logic unlinks the target file during stop or error handling. The file does not include any user-facing disclosure or warning that partial downloads can be automatically deleted, which is a destructive behavior affecting local data.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This code performs HTTP/HTTPS requests and writes downloaded content to disk, but the file contains no confirmation prompt, print/log statement, or explanatory comment/docstring warning users about these actions. Because it handles external downloads and persistent file creation as core behaviors, users invoking surrounding code may not receive explicit notice from this implementation alone.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This JavaScript file contains natural-language comments and user-visible status messages exclusively in Chinese, indicating the skill assumes a specific language/locale. The policy requires flagging language or locale constraints when the skill does not offer user choice or document a justified regional scope.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

The manifest describes transcribing videos from local files or social-media links and producing polished text outputs. This helper exposes a generic download(url, path) function that can fetch any URL to an arbitrary filesystem path, which is a broader capability than is justified by the stated purpose in this file alone.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

This code constructs and sends an HTTPS POST request containing arbitrary JSON payload data plus a TOKEN header, which may include user or system data. In this file, there is no confirmation prompt, user-facing notice, or explanatory comment/docstring disclosing that data and credentials are transmitted to a remote service.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The user-facing strings in this file are entirely in Chinese, including warnings and operational guidance, with no indication that the user can choose another language. This creates a language-policy issue because the skill imposes a specific locale on all users rather than offering or documenting an opt-in choice.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The comment makes strong privacy and data-handling promises (no disclosure, no secondary use, automatic deletion after processing), but this function only performs an upload to a presigned URL and contains no technical controls to enforce deletion, retention limits, or use restrictions. In a video-to-transcript skill, users are likely uploading sensitive personal or business media, so this mismatch can create a real security/privacy issue by causing overtrust in protections that are not actually implemented.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code embeds all visible CLI descriptions, examples, and status/error messages in Chinese, which effectively forces a specific language for interaction. The file does not indicate that the skill is region-specific or provide any opt-in/selection for language, matching the locale policy violation criteria.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest declares the skill as guaikei-video-voice-script, but the documentation's convention section says the canonical name, package directory, package.json name, and H1 should be guaikei-video-to-doc. This is a direct contradiction in the file's own intent/documentation rather than a mere omission.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
99% confidence
Finding

The actual frontmatter at the top of the file uses name: guaikei-video-voice-script, but section 10.2 claims the skill's name is guaikei-video-to-doc. This is an explicit documentation contradiction about the skill's identity and contract.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

Section 10.2 states that tags should be ≤30 and even claims the skill has 30 个, but the frontmatter list runs from L019 through L048, which is 30+ entries. This is a concrete contradiction between the file's stated metadata contract and its actual content.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

Earlier sections repeatedly state that --id last and local task reuse are retained for 24 hours. The FAQ then says users can return within one hour using a task ID, which conflicts with the broader 24-hour retention guidance and creates inconsistent operator expectations.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The contact section is presented only in Chinese and assumes a Chinese-language support channel, which may conflict with a language/locale neutrality policy when the skill otherwise targets broad usage across platforms and languages. There is no opt-in or alternative locale/language path for non-Chinese users.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.