Back to skill

Security audit

智谱免费图片与视频生成

Security checks for vulnerabilities and agentic risk

Overview

This skill coherently provides Zhipu image and video generation through disclosed local Node scripts and API-key-backed calls, with some operational risks users should understand.

Install only if you intend to use Zhipu's external image/video APIs and are comfortable providing a Zhipu API key. Avoid sending sensitive prompts or private image URLs, and be cautious with long video wait settings because malformed polling parameters could keep the process running and generate excess API traffic.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/wait_for_video.js:6
Finding

Unbounded Video Polling Can Cause Resource Exhaustion

Content
View full analysis

Vulnerability Details

File Location: scripts/wait_for_video.js, lines 6–19
Vulnerability Type: Improper validation and bounding of user-controlled polling parameters
Risk Level: Medium

Vulnerable Code

js
const maxWait = Number(input.max_wait_time || 300000);
const poll = Number(input.poll_interval || 5000);
const start = Date.now();

while (Date.now() - start < maxWait) {
  const data = await getJson(`/paas/v4/async-result/${input.task_id}`);
  const status = data.task_status;
  if (status === 'SUCCESS') {
    return print({ success: true, action: 'wait_for_video', task_id: input.task_id, data });
  }
  if (status === 'FAIL') {
    throw new Error(data.error || 'video generation failed');
  }
  await sleep(poll);
}

Technical Analysis

The script converts the attacker-controlled max_wait_time and poll_interval values with Number() but does not verify that they are finite, positive, or within safe limits.

A value such as "Infinity" for max_wait_time makes the loop's time condition remain true indefinitely. A negative value for poll_interval is passed to setTimeout through sleep() and effectively results in little or no delay. Combined, these values can cause the process to issue authenticated polling requests continuously until the task succeeds, fails, or the process is externally terminated.

The script also lacks a maximum request count, cancellation mechanism, and network request timeout, compounding the availability risk.

Attack Path

  1. An attacker or untrusted caller invokes wait_for_video.js with an existing or long-running video task ID.

  2. The caller supplies a payload such as:

    json
    {
      "task_id": "valid-pending-task-id",
      "max_wait_time": "Infinity",
      "poll_interval": -1
    }
    
  3. Number("Infinity") evaluates to positive infinity, while the negative polling interval produces an effectively immediate timer.

  4. The loop repeatedly calls the Zhipu asynchronous-result e ...[truncated 768 chars]

Remediation
View remediation

Remediation Suggestions

Validate both parameters before entering the polling loop:

  1. Require numeric, finite, integer values using Number.isFinite() and Number.isInteger().
  2. Reject non-positive values rather than silently coercing them.
  3. Enforce explicit limits, for example:
    • poll_interval: 1,000–30,000 milliseconds.
    • max_wait_time: 1,000–900,000 milliseconds.
  4. Add a maximum polling-attempt count independent of elapsed time.
  5. Add a timeout to each HTTPS request so a stalled connection cannot hold the process indefinitely.
  6. Support cancellation through an AbortController or equivalent mechanism.
  7. Consider applying exponential backoff with a maximum delay.

Example validation:

js
const maxWait = Number(input.max_wait_time ?? 300000);
const poll = Number(input.poll_interval ?? 5000);

if (!Number.isFinite(maxWait) || !Number.isInteger(maxWait) ||
    maxWait < 1000 || maxWait > 900000) {
  throw new Error('max_wait_time must be an integer between 1000 and 900000');
}

if (!Number.isFinite(poll) || !Number.isInteger(poll) ||
    poll < 1000 || poll > 30000) {
  throw new Error('poll_interval must be an integer between 1000 and 30000');
}

const maxAttempts = Math.ceil(maxWait / poll);
for (let attempt = 0; attempt < maxAttempts; attempt++) {
  // Query status and stop on success or failure.
  await sleep(poll);
}
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (5)

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 55)May include surrounding context.

md
- `scripts/query_video_result.js` - 查询视频任务结果

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 56)May include surrounding context.

md
- `scripts/wait_for_video.js` - 等待视频生成完成

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill documents use of environment-provided API keys and executable Node scripts, but it does not declare any explicit tool scope such as permissions or allowed-tools. That creates an authorization gap where the runtime may grant broader-than-intended access to env-backed secrets or execution capabilities, making accidental secret exposure or overly broad tool use more likely. In this context, the risk is real because the skill is specifically designed to invoke scripts that consume API credentials.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrases include broad everyday requests such as '帮我生成一张图' or '生成一个短视频', which can cause the skill to activate in situations where the user did not specifically intend to use this capability. Mis-triggering is dangerous because it can route unrelated requests into a workflow that invokes scripts and external APIs, potentially consuming credits or using API-backed capabilities unnecessarily. The skill context increases plausibility because it is built for direct execution of image/video generation workflows.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code posts the user's prompt and optional image_url to an external video-generation endpoint, which is a data-transmitting network operation. There is no confirmation prompt, warning message, or explanatory comment/docstring in this file disclosing that user content will be sent to a remote service.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.