T09 · Insecure Skill Coding Practices
- Location
script/glm_understanding.py:27- Finding
Prompt Injection Through Untrusted OCR-Derived Document Content
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This is a cloud-based document OCR and analysis skill that uploads user-selected documents to Zhipu GLM services and writes local reports, which fits its stated purpose.
Install only if you are comfortable sending the documents, tables, captions, and cropped images you process to Zhipu GLM services using your API key. Avoid confidential or regulated documents unless your policy allows that provider, use a protected output directory, and treat generated analysis as model output that may be influenced by malicious text inside the source document.
script/glm_understanding.py:27Prompt Injection Through Untrusted OCR-Derived Document Content
The code base64-encodes the entire local file and submits it to client.layout_parsing.create, which sends document contents to an external OCR provider. If the input contains sensitive business, personal, or regulated data, this can result in unauthorized disclosure outside the local trust boundary.
This code reads a local image file, base64-encodes it, and uploads it together with up to 6000 characters of document context to a remote multimodal model. That creates a direct exfiltration path for potentially sensitive local files and surrounding document contents, with no policy enforcement, user approval, or restriction on what images may be sent.
The skill explicitly routes document text, tables, and cropped images to external Zhipu-hosted models, but the user-facing description does not clearly disclose that potentially sensitive document contents leave the local environment. This creates a real privacy and data-governance risk because users may submit confidential PDFs or images without informed consent or understanding of third-party processing.
The skill documentation describes sending table text, full Markdown context, and Base64-encoded cropped images to external GLM services, but it does not clearly warn users that potentially sensitive document content leaves the local environment. This creates a real privacy and data-governance risk because users may process confidential PDFs or images without understanding that third-party API transmission occurs.
The script is explicitly configured to use a remote Zhipu OCR client with an API key from the environment, establishing an outbound data path for user-supplied files. In the absence of clear consent, disclosure, or data-handling constraints, this creates a real confidentiality risk because local documents and images may be sent to a third-party service.
The script transmits image content to an external OCR service without any user-facing notice or confirmation at the point of use. That omission increases the chance that operators unknowingly upload confidential material, making accidental data leakage more likely in real deployments.
The script initializes a remote Zhipu AI client using an API key from the environment and later sends document-derived content to that external service. In a security context, this is a real data-exposure risk because the skill provides no consent flow, no data classification checks, and no boundary limiting what local content may be transmitted.
The text-model API call transmits document context to an external service, but the script only emits generic progress logs and gives no meaningful warning that local content is leaving the environment. This is dangerous because operators may process confidential documents under the false assumption that analysis is local.
The vision request silently uploads raw local image contents to an external API without explicit notice at the time of processing. Because images can contain sensitive screenshots, figures, identifiers, or embedded document fragments, the lack of warning materially increases the chance of accidental disclosure.
The skill states that outputs include cropped image files and JSON reports in an output directory, but it does not clearly warn that derived artifacts are written to disk. This is a genuine information-exposure risk because extracted figures and analysis outputs may contain sensitive content and remain on local or shared storage after processing.
No suspicious patterns detected.