Back to skill

Security audit

Zai Vision

Security checks for vulnerabilities and agentic risk

Overview

This skill does what it claims: it analyzes user-selected images or videos through Z.AI, with normal privacy and dependency risks to consider.

Install this only if you are comfortable sending selected images or videos to Z.AI for processing. Avoid using it on confidential screenshots, credentials, regulated data, or private videos unless your organization permits that. Prefer installing `zai-sdk` in a virtual environment, pinning a reviewed version, and keeping `ZAI_API_KEY` in the environment rather than source files.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:18
Finding

Unpinned Third-Party SDK Installation Creates a Supply-Chain Risk

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (10)

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

The description claims a general-purpose vision skill covering both image and video understanding, including OCR and analysis of screenshots, UI designs, photos, diagrams, and charts. However, the provided code chunk only implements video analysis from a local file path. It does not include any image input handling, image MIME routing, OCR-specific logic, or separate support for the listed image-centric use cases. The core behavior—sending media to Z.AI's vision model for analysis—is aligned in spirit with the video portion of the description, but the implemented capability is materially narrower than the declared purpose. There are no obvious undeclared dangerous capabilities beyond reading a local file and calling the API, but the description does not accurately represent the full breadth of what this code chunk actually supports.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The code substantially matches the image-analysis portion of the description: it uses Z.AI's GLM-4.6V vision model, accepts an image plus prompt, and can support OCR/UI/chart/diagram analysis indirectly through prompting. However, the description explicitly includes video understanding and video scene description, while the supplied code has no video input support, no video preprocessing, and no API usage for video content. Because video capability is a declared part of the skill's purpose but is absent from the actual implementation, this is a material description-behavior mismatch.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/video_analyze.py (reported line 76)May include surrounding context.

python
Returns:
        Analysis result text or JSON object
    """
    # Get API key from environment or use default
    api_key = os.getenv("ZAI_API_KEY")

    if not api_key:

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/vision_analyze.py (reported line 57)May include surrounding context.

python
Returns:
        Analysis result text or JSON object
    """
    # Get API key from environment or use default
    api_key = os.getenv("ZAI_API_KEY")

    if not api_key:

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill repeatedly instructs users to analyze local images and videos via the Z.AI Vision API but does not clearly warn that those files may be uploaded to a third-party external service. This can lead to inadvertent disclosure of sensitive screenshots, documents, diagrams, or videos containing secrets, personal data, or internal business information.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The documentation instructs users to analyze local images and videos with a third-party vision API but does not clearly warn that the referenced files will be uploaded off-host for remote processing. In a vision-analysis skill, users may pass screenshots, UI captures, diagrams, or videos that contain secrets, personal data, or internal system details, so the omission creates a real risk of unintended data exfiltration.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The Python example initializes the client with a literal API key placeholder in code, which normalizes hardcoding credentials and may lead users to paste real secrets into source files, notebooks, or shared snippets. Hardcoded API keys are commonly leaked through version control, logs, screenshots, or copied examples, enabling unauthorized API use and possible billing abuse or access to associated data.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script base64-encodes the supplied video and sends it to a third-party API for processing, but it does not provide an explicit user-facing warning or consent checkpoint at the point of transmission. Because videos may contain sensitive visual data, this can lead to unintended disclosure of confidential or personal information in normal use.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The script base64-encodes the supplied image and sends it to a third-party API, but it does not provide any explicit warning, confirmation, or consent mechanism before transmitting potentially sensitive local image content off-host. In a vision-analysis skill, this behavior is expected functionally, but it still creates a real privacy and data-handling risk because screenshots, photos, diagrams, or videos may contain credentials, personal data, or proprietary information.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.