Back to skill

Security audit

GLM-OCR

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent OCR skill that uses a fixed GLM-OCR API endpoint, with ordinary privacy and API-key handling caveats but no evidence of hidden or deceptive behavior.

Install only if you are comfortable sending selected images, PDFs, or URLs to the GLM-OCR service. Treat the API key as sensitive: prefer environment variables or a protected .env file, avoid putting real keys directly in shell history, and consider pinning dependencies before use in stricter environments.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/config_setup.py:65
Finding

API Credential Persisted Without Enforced Restrictive File Permissions

Content
View full analysis
None: """ Write environment variables to .env file Args: env_vars: Dictionary of environment variables to write skill_root: Skill root directory path """ env_path = _get_env_path(skill_root) with open(env_path, "w", encoding="utf-8") as f: # Write header f.write("# GLM-OCR API Configuration\n") f.write("# Auto-generated by config_setup.py\n") f.write("# DO NOT commit this file to version control\n\n") # Write environment variables for key, value in env_vars.items(): f.write(f"{key}={value}\n") ``` ### Technical Analysis The setup process writes `ZHIPU_API_KEY` in plaintext to the project-level `.env` file. The code neither creates the file with an explicit owner-only mode nor corrects the permissions of an existing file. Consequently, the resulting permissions depend on the process umask and any permissions already assigned to the file. Under a permissive umask, the file may be readable by other local users or processes. Writing a warning inside the file does not enforce access control. The documented setup form also accepts the API key as a command-line argument: ```bash python scripts/config_setup.py setup --api-key YOUR_KEY ``` This may leave the credential in shell history and can expose it through process-argument inspection while the command is running. This exposure is related to credential provisioning, although the permission defect is specifically located in `_write_env_file()`. ### Attack Path 1. A user runs the documented setup command and supplies a valid GLM-OCR API key. 2. `config_setup.py` writes the key in plaintext to `.env` ...[truncated 1082 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
scripts/requirements.txt:1
Finding

Unbounded and Integrity-Unlocked Third-Party Dependency

Content
View full analysis
=2.31.0 ``` ### Technical Analysis The requirement specifies only a minimum version of `requests`. It permits the package resolver to install any later release that satisfies the constraint. The project also provides no lock file or package hashes. As a result, two installations performed at different times can resolve to different dependency versions without any change to the audited Skill. This weakens reproducibility and allows future, unreviewed dependency releases to enter the execution environment automatically. No evidence was found that the named package is currently malicious, misspelled, or retrieved from an explicitly unsafe package source. The risk arises from the open-ended version range and absence of integrity locking rather than from a confirmed compromise of `requests`. ### Attack Path 1. A user or deployment system installs dependencies from `scripts/requirements.txt`. 2. The package resolver selects the latest release satisfying `requests>=2.31.0`. 3. A future compromised, malicious, or incompatible release is available from the configured package index. 4. That unreviewed release is installed because no exact version or expected artifact hash is enforced. 5. The dependency executes when imported by `scripts/glm_ocr_cli.py`. 6. A compromised dependency could access the process environment, API credential, local document data, and outbound OCR requests under the privileges of the invoking user. This attack path depends on a future dependency or package-distribution compromise; the audit found no evidence of a current malicious release. ### Impact Assessment A compromised dependency would execute with the same operating-system privileges as the OCR script. Within that scope, it could potentially: ...[truncated 530 chars]
Remediation
View remediation
``` 2. Generate and maintain a lock file that also pins transitive dependencies. 3. Record cryptographic hashes for approved distribution artifacts and install with hash verification, for example through pip's `--require-hashes` mode. 4. Retrieve packages only from a trusted, explicitly configured package index. 5. Review dependency updates before changing the lock file, including release notes and known-vulnerability advisories. 6. Use automated dependency scanning while retaining manual approval for updates that handle credentials, document contents, or network traffic. 7. Rebuild and test the OCR workflow after every dependency update to detect behavior or compatibility changes. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (23)

Tainted flow: 'timeout' from os.getenv (line 181, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/glm_ocr_cli.py (reported line 184)May include surrounding context.

python
timeout = float(os.getenv("GLM_OCR_TIMEOUT", str(DEFAULT_TIMEOUT)))

    try:
        resp = requests.post(api_url, json=payload, headers=headers, timeout=timeout)
    except requests.Timeout:
        raise RuntimeError(f"API request timed out after {timeout}s")
    except requests.RequestException as e:

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill is presented as a simple OCR tool, but the documentation also exposes configuration-management behavior such as setting up API keys, reading and writing environment files, and displaying masked secret-related configuration state. This mismatch weakens user consent and reviewability because users may invoke the skill for OCR without realizing it also performs local configuration operations beyond the primary task.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config_setup.py (reported line 4)May include surrounding context.

python
#!/usr/bin/env python3
"""
Environment Variable Configuration Setup
Helps users set up their .env file for GLM-OCR skill
"""

import sys

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config_setup.py (reported line 20)May include surrounding context.

python
#!/usr/bin/env python3
"""
Environment Variable Configuration Setup
Helps users set up their .env file for GLM-OCR skill
"""

import sys

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config_setup.py (reported line 27)May include surrounding context.

python
#!/usr/bin/env python3
"""
Environment Variable Configuration Setup
Helps users set up their .env file for GLM-OCR skill
"""

import sys

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config_setup.py (reported line 33)May include surrounding context.

python
#!/usr/bin/env python3
"""
Environment Variable Configuration Setup
Helps users set up their .env file for GLM-OCR skill
"""

import sys

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config_setup.py (reported line 60)May include surrounding context.

python
#!/usr/bin/env python3
"""
Environment Variable Configuration Setup
Helps users set up their .env file for GLM-OCR skill
"""

import sys

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config_setup.py (reported line 164)May include surrounding context.

python
#!/usr/bin/env python3
"""
Environment Variable Configuration Setup
Helps users set up their .env file for GLM-OCR skill
"""

import sys

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config_setup.py (reported line 170)May include surrounding context.

python
#!/usr/bin/env python3
"""
Environment Variable Configuration Setup
Helps users set up their .env file for GLM-OCR skill
"""

import sys

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config_setup.py (reported line 252)May include surrounding context.

python
#!/usr/bin/env python3
"""
Environment Variable Configuration Setup
Helps users set up their .env file for GLM-OCR skill
"""

import sys

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/glm_ocr_cli.py (reported line 56)May include surrounding context.

python
#!/usr/bin/env python3
"""
Environment Variable Configuration Setup
Helps users set up their .env file for GLM-OCR skill
"""

import sys

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config_setup.py (reported line 23)May include surrounding context.

python
"""Get the path to the .env file"""
    if skill_root is None:
        skill_root = _get_skill_root()
    return skill_root / ".env"


def _env_exists(skill_root: Optional[Path] = None) -> bool:

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config_setup.py (reported line 264)May include surrounding context.

python
"""Get the path to the .env file"""
    if skill_root is None:
        skill_root = _get_skill_root()
    return skill_root / ".env"


def _env_exists(skill_root: Optional[Path] = None) -> bool:

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/glm_ocr_cli.py (reported line 61)May include surrounding context.

python
"""Get the path to the .env file"""
    if skill_root is None:
        skill_root = _get_skill_root()
    return skill_root / ".env"


def _env_exists(skill_root: Optional[Path] = None) -> bool:

Self-Modification

High
Category
Rogue Agent
Confidence
85% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · scripts/config_setup.py (reported line 63)May include surrounding context.

python
Write environment variables to .env file

    Args:
        env_vars: Dictionary of environment variables to write
        skill_root: Skill root directory path
    """
    env_path = _get_env_path(skill_root)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config_setup.py (reported line 265)May include surrounding context.

python
with open(gitignore_path, "r") as f:
                gitignore_content = f.read()
            if ".env" not in gitignore_content:
                print("Tip: Add '.env' to your .gitignore to keep API keys secure")
        else:
            print("Tip: Create a .gitignore file with '.env' to keep API keys secure")

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config_setup.py (reported line 267)May include surrounding context.

python
with open(gitignore_path, "r") as f:
                gitignore_content = f.read()
            if ".env" not in gitignore_content:
                print("Tip: Add '.env' to your .gitignore to keep API keys secure")
        else:
            print("Tip: Create a .gitignore file with '.env' to keep API keys secure")

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill declares no explicit tool scope even though its documented behavior requires environment access, local file read/write, and network access. Without a restrictive permissions declaration, an agent platform may grant broader capabilities than users expect, increasing the risk of unintended data access or exfiltration when processing local files and URLs.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The activation text is broad enough to trigger on general image-processing requests, not just explicit OCR tasks. That can cause the agent to send files or URLs to an external OCR service in situations where the user did not intend remote processing, creating privacy and consent risks.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill does not clearly warn that local files or remote URLs provided by the user are transmitted to a third-party OCR API for processing. In an OCR context, uploaded images and PDFs may contain sensitive personal, financial, legal, or corporate information, so lack of disclosure undermines informed consent and can lead to unintended external data exposure.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
93% confidence
Finding

This skill transmits user-supplied content to an external OCR service, including full local file contents after base64 encoding or remote URLs for server-side fetching. In the context of an OCR skill this is expected behavior, but it is still security-relevant because sensitive documents may be sent off-host to a third-party API.

Content

Scanner excerpt · scripts/glm_ocr_cli.py (reported line 184)May include surrounding context.

python
timeout = float(os.getenv("GLM_OCR_TIMEOUT", str(DEFAULT_TIMEOUT)))

    try:
        resp = requests.post(api_url, json=payload, headers=headers, timeout=timeout)
    except requests.Timeout:
        raise RuntimeError(f"API request timed out after {timeout}s")
    except requests.RequestException as e:

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
95% confidence
Finding

The dependency is specified as requests>=2.31.0, which allows any future major or minor release to be installed without review. This weakens reproducibility and can introduce vulnerable or breaking versions through the supply chain, even if the currently resolved version is safe.

Content

Scanner excerpt · scripts/requirements.txt (reported line 1)May include surrounding context.

text
requests>=2.31.0

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
86% confidence
Finding

Because the manifest does not pin the requests version, it is impossible to verify from this file whether deployment will use a release affected by known advisories. In an OCR skill that likely makes outbound HTTP requests to an API, using an affected requests version could expose credentials, weaken TLS verification behavior, or otherwise impact network security depending on runtime resolution.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.