Back to skill

Security audit

Guaikei Video2text 1.0.1

Security checks for vulnerabilities and agentic risk

Overview

This skill is a coherent video-to-text cloud integration, but users should understand that provided media is sent to GuaiKei and may incur usage charges.

Install only if you are comfortable sending the specific video or public media URL you provide to GuaiKei's cloud service under your API token. Avoid using it for sensitive internal recordings unless GuaiKei's privacy, retention, and billing terms are acceptable to you.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (4)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The frontmatter description advertises local file and public URL support plus cloud transcription, but it does not clearly warn up front that user-provided video content will be uploaded to an external cloud service for processing and may incur charges. This can cause unintended disclosure of sensitive local media or private information embedded in URLs because users may trigger the skill before understanding the data-transfer boundary.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The trigger section tells the agent when to call the skill, but it omits a nearby notice that invocation sends user content to an external service and can generate billable usage. In an agent setting, this increases the risk of surprise exfiltration and unexpected charges because the tool may be auto-selected based only on user intent phrases.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · SKILL.md (reported line 158)May include surrounding context.

md
## 7. 🗣️ 自然语言 → 命令(照这张表转)

| 用户说的话                                           | 就执行这条命令                                                                                              |
| ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 提取 https://example.com/video.mp4 里的文字          | `node scripts/video2text/index.js --file "https://example.com/video.mp4"`                                   |
| 总结这个视频的核心观点 https://example.com/video.mp4 | `node scripts/video2text/index.js --file "https://example.com/video.mp4" --prompt "总结这个视频的核心观点"` |

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The documentation consistently instructs in Chinese and includes default prompt templates oriented to Chinese-language outputs and Chinese platforms, without an explicit user language-choice statement in the general skill instructions. This can be read as a locale/language default that is not explicitly opt-in.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.