Back to skill

Security audit

GLM-V-Caption

Security checks for vulnerabilities and agentic risk

Overview

This appears to be a straightforward Zhipu captioning skill, but users should know submitted media is sent to Zhipu and the docs have a few accuracy issues.

Install only if you are comfortable sending selected images, video URLs, document URLs, prompts, and your Zhipu API usage to Zhipu's service. Avoid confidential, regulated, or secret-containing media unless your organization permits that provider, and be aware that videos and documents are URL-only despite some path-oriented wording in the docs.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Rogue AgentSelf-Modification, Session Persistence
Findings (14)

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The primary purpose matches: this is indeed a caption/description tool for multimodal inputs using Zhipu models. There are no obvious undeclared dangerous capabilities beyond making outbound API requests to Zhipu. However, the declared input support is materially broader than the implementation for videos and documents. The description says the skill supports URLs, local paths, and base64 (images only) for images, videos, and files, but the code explicitly enforces that videos and files must be URLs and returns errors for local paths. That is a meaningful behavior-description mismatch. There is also minor model-description inconsistency: the file presents itself as GLM-4.6V-oriented, but the actual default model constant is glm-5v-turbo. This is secondary, but it adds to representational inconsistency.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 71)May include surrounding context.

export ZHIPU_API_KEY="你的密钥"

text

3. **.env file / .env 文件:** Create `.env` in this skill directory:

ZHIPU_API_KEY=你的密钥

text

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
88% confidence
Finding

The instruction to always return the model's full raw output exactly as received can cause unreviewed disclosure of sensitive data, harmful content, prompt artifacts, or provider-returned details that should be filtered or minimized before presentation. In a multimodal/document skill, this is more dangerous because uploaded files may contain private or regulated information that the downstream model could echo verbatim.

Content

Scanner excerpt · SKILL.md (reported line 84)May include surrounding context.

md
4. **IF API fails** — Display the error message and STOP immediately
5. **NO fallback methods** — Do NOT attempt captioning any other way

### 📋 Output Display Rules (MANDATORY)

After running the script, **you must show the full raw output to the user exactly as returned**. Do not summarize, truncate, or only say "generated". Users need the original model output to evaluate quality.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill declares no explicit tool scope even though it clearly relies on environment variables, network access, and file writing. Missing permission boundaries increases the blast radius if the skill is invoked in a broader agent runtime, because the agent may grant more capability than users expect.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill processes images, videos, documents, and URLs via an external ZhiPu API but does not prominently warn that user-provided content and possibly sensitive document contents are transmitted off-platform. This creates a real privacy and data-handling risk, especially for confidential files, internal URLs, or regulated data.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · SKILL.md (reported line 43)May include surrounding context.

md
## Resource Links

| Resource        | Link                                                                                                                              |
| --------------- | --------------------------------------------------------------------------------------------------------------------------------- |
| **Get API Key** | [https://bigmodel.cn/usercenter/proj-mgmt/apikeys](https://bigmodel.cn/usercenter/proj-mgmt/apikeys)                              |
| **API Docs**    | [Chat Completions / 对话补全](https://docs.bigmodel.cn/api-reference/%E6%A8%A1%E5%9E%8B-api/%E5%AF%B9%E8%AF%9D%E8%A1%A5%E5%85%A8) |

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 55)May include surrounding context.

md
This script reads the key from the `ZHIPU_API_KEY` environment variable and shares it with other Zhipu skills.
脚本通过 `ZHIPU_API_KEY` 环境变量获取密钥,与其他智谱技能共用同一个 key。

**Get Key / 获取 Key:** Visit [Zhipu Open Platform API Keys / 智谱开放平台 API Keys](https://bigmodel.cn/usercenter/proj-mgmt/apikeys) to create or copy your key.

**Setup options / 配置方式(任选一种):**

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 55)May include surrounding context.

md
This script reads the key from the `ZHIPU_API_KEY` environment variable and shares it with other Zhipu skills.
脚本通过 `ZHIPU_API_KEY` 环境变量获取密钥,与其他智谱技能共用同一个 key。

**Get Key / 获取 Key:** Visit [Zhipu Open Platform API Keys / 智谱开放平台 API Keys](https://bigmodel.cn/usercenter/proj-mgmt/apikeys) to create or copy your key.

**Setup options / 配置方式(任选一种):**

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · SKILL.md (reported line 145)May include surrounding context.

python {baseDir}/scripts/glmv_caption.py (--images IMG [IMG...] | --videos VID [VID...] | --files FILE [FILE...]) [OPTIONS]

text

| Parameter             | Required | Description                                                                                  |
| --------------------- | -------- | -------------------------------------------------------------------------------------------- |
| `--images`, `-i`      | One of   | Image paths or URLs (supports multiple, base64 OK)                                           |
| `--videos`, `-v`      | One of   | Video paths or URLs (supports multiple, mp4/mkv/mov)                                         |

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The default prompt is hard-coded in Chinese ("请详细描述这张图片的内容"), which imposes a specific language preference on users by default. This is a natural-language locale policy concern because the tool does not offer a language choice or document that the language constraint is intentional and region-specific.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The caption() docstring says videos and files may be provided as local paths, and the CLI help repeats that claim for both input types. However, the implementation explicitly returns errors unless every video and file input is an HTTP(S) URL, so the documentation contradicts actual behavior rather than merely omitting details.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
93% confidence
Finding

The skill sends user-supplied images, file URLs, video URLs, prompts, and potentially local image contents encoded as data URLs to an external third-party API. In an agent context, this is a real data exfiltration boundary: sensitive local images or confidential user content may be transmitted off-system without explicit consent, and remote URLs may cause the provider to retrieve additional sensitive resources.

Content

Scanner excerpt · scripts/glmv_caption.py (reported line 305)May include surrounding context.

python
if thinking:
        payload["thinking"] = {"type": "enabled"}

    response = requests.post(
        API_BASE_URL, headers=headers, json=payload, stream=stream, timeout=120
    )

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documentation at L039 states that videos and files only support URLs and that local paths are not supported for those types. However, the CLI parameter descriptions at L148-L149 say videos and files accept "paths or URLs," which directly conflicts with the stated restrictions and could mislead an agent or user about actual intended behavior.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The module docstring presents the tool as a GLM-4.6V caption tool, yet the actual default model constant is set to 'glm-5v-turbo'. That is an active mismatch between stated intent and effective runtime behavior for users who rely on defaults.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.