Back to skill

Security audit

pdf-ocr-layout

Security checks for vulnerabilities and agentic risk

Overview

This is a cloud-based document OCR and analysis skill that uploads user-selected documents to Zhipu GLM services and writes local reports, which fits its stated purpose.

Install only if you are comfortable sending the documents, tables, captions, and cropped images you process to Zhipu GLM services using your API key. Avoid confidential or regulated documents unless your policy allows that provider, use a protected output directory, and treat generated analysis as model output that may be influenced by malicious text inside the source document.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
script/glm_understanding.py:27
Finding

Prompt Injection Through Untrusted OCR-Derived Document Content

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (10)

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The code base64-encodes the entire local file and submits it to client.layout_parsing.create, which sends document contents to an external OCR provider. If the input contains sensitive business, personal, or regulated data, this can result in unauthorized disclosure outside the local trust boundary.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

This code reads a local image file, base64-encodes it, and uploads it together with up to 6000 characters of document context to a remote multimodal model. That creates a direct exfiltration path for potentially sensitive local files and surrounding document contents, with no policy enforcement, user approval, or restriction on what images may be sent.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill explicitly routes document text, tables, and cropped images to external Zhipu-hosted models, but the user-facing description does not clearly disclose that potentially sensitive document contents leave the local environment. This creates a real privacy and data-governance risk because users may submit confidential PDFs or images without informed consent or understanding of third-party processing.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill documentation describes sending table text, full Markdown context, and Base64-encoded cropped images to external GLM services, but it does not clearly warn users that potentially sensitive document content leaves the local environment. This creates a real privacy and data-governance risk because users may process confidential PDFs or images without understanding that third-party API transmission occurs.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The script is explicitly configured to use a remote Zhipu OCR client with an API key from the environment, establishing an outbound data path for user-supplied files. In the absence of clear consent, disclosure, or data-handling constraints, this creates a real confidentiality risk because local documents and images may be sent to a third-party service.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script transmits image content to an external OCR service without any user-facing notice or confirmation at the point of use. That omission increases the chance that operators unknowingly upload confidential material, making accidental data leakage more likely in real deployments.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script initializes a remote Zhipu AI client using an API key from the environment and later sends document-derived content to that external service. In a security context, this is a real data-exposure risk because the skill provides no consent flow, no data classification checks, and no boundary limiting what local content may be transmitted.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The text-model API call transmits document context to an external service, but the script only emits generic progress logs and gives no meaningful warning that local content is leaving the environment. This is dangerous because operators may process confidential documents under the false assumption that analysis is local.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The vision request silently uploads raw local image contents to an external API without explicit notice at the time of processing. Because images can contain sensitive screenshots, figures, identifiers, or embedded document fragments, the lack of warning materially increases the chance of accidental disclosure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill states that outputs include cropped image files and JSON reports in an output directory, but it does not clearly warn that derived artifacts are written to disk. This is a genuine information-exposure risk because extracted figures and analysis outputs may contain sensitive content and remain on local or shared storage after processing.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.