T08 · Insecure Dependencies
- Location
SKILL.md:18- Finding
Unpinned Third-Party SDK Installation Creates a Supply-Chain Risk
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill does what it claims: it analyzes user-selected images or videos through Z.AI, with normal privacy and dependency risks to consider.
Install this only if you are comfortable sending selected images or videos to Z.AI for processing. Avoid using it on confidential screenshots, credentials, regulated data, or private videos unless your organization permits that. Prefer installing `zai-sdk` in a virtual environment, pinning a reviewed version, and keeping `ZAI_API_KEY` in the environment rather than source files.
SKILL.md:18Unpinned Third-Party SDK Installation Creates a Supply-Chain Risk
The description claims a general-purpose vision skill covering both image and video understanding, including OCR and analysis of screenshots, UI designs, photos, diagrams, and charts. However, the provided code chunk only implements video analysis from a local file path. It does not include any image input handling, image MIME routing, OCR-specific logic, or separate support for the listed image-centric use cases. The core behavior—sending media to Z.AI's vision model for analysis—is aligned in spirit with the video portion of the description, but the implemented capability is materially narrower than the declared purpose. There are no obvious undeclared dangerous capabilities beyond reading a local file and calling the API, but the description does not accurately represent the full breadth of what this code chunk actually supports.
The code substantially matches the image-analysis portion of the description: it uses Z.AI's GLM-4.6V vision model, accepts an image plus prompt, and can support OCR/UI/chart/diagram analysis indirectly through prompting. However, the description explicitly includes video understanding and video scene description, while the supplied code has no video input support, no video preprocessing, and no API usage for video content. Because video capability is a declared part of the skill's purpose but is absent from the actual implementation, this is a material description-behavior mismatch.
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
Returns:
Analysis result text or JSON object
"""
# Get API key from environment or use default
api_key = os.getenv("ZAI_API_KEY")
if not api_key:
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
Returns:
Analysis result text or JSON object
"""
# Get API key from environment or use default
api_key = os.getenv("ZAI_API_KEY")
if not api_key:
Without declared permissions the skill's intent is opaque and cannot be validated.
The skill repeatedly instructs users to analyze local images and videos via the Z.AI Vision API but does not clearly warn that those files may be uploaded to a third-party external service. This can lead to inadvertent disclosure of sensitive screenshots, documents, diagrams, or videos containing secrets, personal data, or internal business information.
The documentation instructs users to analyze local images and videos with a third-party vision API but does not clearly warn that the referenced files will be uploaded off-host for remote processing. In a vision-analysis skill, users may pass screenshots, UI captures, diagrams, or videos that contain secrets, personal data, or internal system details, so the omission creates a real risk of unintended data exfiltration.
The Python example initializes the client with a literal API key placeholder in code, which normalizes hardcoding credentials and may lead users to paste real secrets into source files, notebooks, or shared snippets. Hardcoded API keys are commonly leaked through version control, logs, screenshots, or copied examples, enabling unauthorized API use and possible billing abuse or access to associated data.
The script base64-encodes the supplied video and sends it to a third-party API for processing, but it does not provide an explicit user-facing warning or consent checkpoint at the point of transmission. Because videos may contain sensitive visual data, this can lead to unintended disclosure of confidential or personal information in normal use.
The script base64-encodes the supplied image and sends it to a third-party API, but it does not provide any explicit warning, confirmation, or consent mechanism before transmitting potentially sensitive local image content off-host. In a vision-analysis skill, this behavior is expected functionally, but it still creates a real privacy and data-handling risk because screenshots, photos, diagrams, or videos may contain credentials, personal data, or proprietary information.
No suspicious patterns detected.