T09 · Insecure Skill Coding Practices
- Location
scripts/lib.py:81- Finding
Sensitive documents and authentication credentials may be transmitted over plaintext HTTP
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill has a coherent document-parsing purpose, but it may send full documents and credentials over plaintext HTTP and saves raw OCR output to temp files by default.
Install only if you intend to send documents to the configured PaddleOCR/Triton service. Use HTTPS for any non-local endpoint, avoid sending access tokens or Basic Auth over HTTP, and use --stdout or a private output path for sensitive documents so raw OCR results are not left in the system temp directory.
scripts/lib.py:81Sensitive documents and authentication credentials may be transmitted over plaintext HTTP
scripts/vl_caller.py:47Complete OCR results are stored in a shared temporary hierarchy without explicit access restrictions or cleanup
The declared description centers on OCR-based document parsing and structured conversion of PDFs/images into Markdown and JSON. The actual code only validates image file extensions, opens images with Pillow, converts color modes if needed, saves them with specified quality, and iteratively resizes them to reduce file size. It does not invoke PaddleOCR, perform OCR, extract document structure, process PDFs, or emit Markdown/JSON. While file optimization could be a supporting preprocessing step in a larger document-parsing system, this code chunk’s primary purpose is materially different from the declared purpose, so this is a clear mismatch.
The declared description promises a document-understanding/OCR pipeline that extracts content from PDFs and images into structured Markdown and JSON. The actual code only validates CLI arguments, parses page ranges, opens a PDF, imports selected pages into a new PDF, and saves the result. It neither performs OCR nor extracts text, layout, or structure, and it does not emit Markdown or JSON. This is a clear material mismatch in primary purpose and capabilities.
The code generally aligns with a document-parsing purpose using PaddleOCR, but the declared description overstates or differs from the demonstrated behavior in important ways. The shown code only serializes parser output to JSON and saves/prints it; there is no visible Markdown generation despite the description claiming conversion into Markdown and JSON files. It also supports fetching documents from a URL and writing results to local storage, capabilities not reflected in the declared permissions/description. While filesystem output is a supporting detail for a parser CLI, remote URL ingestion and the absence of Markdown support make the description not fully accurate for this code chunk.
Referenced artifact was not completely inspected
pip install -r scripts/requirements-optimize.txt
The skill is advertised as local document parsing, but this script requires users to obtain remote API credentials and connect to an external PaddleOCR endpoint. That mismatch can mislead users into exposing documents, metadata, and secrets to a third-party service under the assumption that processing is local-only.
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
2. Open your model's API page and sign in
3. Open your model's Example Code section
4. In Example Code, copy the API URL value
5. In Example Code, copy the Access Token value
Set environment variables:
export PADDLEOCR_DOC_PARSING_API_URL=https://your-api-url.paddleocr.com/layout-parsing
The skill invokes Python scripts, reads local files, writes JSON results to disk, uses environment variables, and sends documents to a configured network endpoint, but it does not declare any explicit tool restrictions such as allowed-tools or permissions. That creates unnecessary ambient authority: an agent or runtime may grant broader shell, filesystem, and network access than users expect, increasing the blast radius if the skill is misused or later modified.
The skill instructs the agent to send user-supplied local files or URLs to a configured API endpoint and to persist raw extracted JSON to disk, but it does not foreground a clear privacy and data-handling warning. Users may unintentionally transmit sensitive documents, credentials embedded in URLs, or regulated data to a remote service or local temp storage without informed consent.
The documentation states that parsed output is automatically written to a unique file in the system temp directory, but does not warn that the output may contain sensitive document text or extracted structured data. Persisting such data by default increases the risk of unintended disclosure through shared temp locations, leftover files, forensic recovery, or other local users/processes accessing sensitive OCR results.
The documentation explicitly describes parsing documents from URLs, which extends the skill's behavior beyond an expected local-only scope and can lead users or integrating agents to fetch attacker-controlled remote content. That expands the trust boundary and introduces SSRF-like, data exfiltration, and unexpected network access risks, especially if consumers rely on the manifest or skill description to determine safety constraints.
The code reads access tokens and basic-auth credentials from environment variables and performs HTTP requests to an API endpoint via httpx. For a skill described as local PaddleOCR parsing, outbound network communication and auth handling are not obviously justified unless the manifest explicitly states dependence on a local/remote inference service.
Local document bytes are base64-encoded and sent over HTTP(S) to a configured endpoint without any explicit user-facing consent or warning at the call site. For a skill presented as local document parsing, this can cause unanticipated disclosure of sensitive document contents to another service, especially if the endpoint is remote, misconfigured, or uses plain HTTP.
The function accepts an arbitrary file_url and forwards it to the backend parsing service, which will then retrieve the remote resource. This creates an SSRF-style capability in the backend and contradicts the 'local document parsing' description, potentially allowing access to internal-only URLs or unintended outbound requests if an attacker controls the URL input.
The manifest describes a local PaddleOCR-based document parsing skill that converts PDFs and images into Markdown and JSON. Declaring an HTTP client dependency suggests network access capability, which is not an obvious requirement for purely local OCR and document conversion.
The manifest presents the skill as operating locally to parse PDFs and document images, but this script's documented purpose is to verify API connectivity and the implementation sends a document URL to a remote PaddleOCR endpoint. That is a meaningful behavioral mismatch because network-based remote processing is not implied by a 'locally' named parsing skill.
The CLI advertises local document parsing but also accepts remote URLs and depends on external API configuration, which changes the trust boundary and can cause users to send sensitive documents or trigger network access they did not expect. In an agent/skill context, undisclosed remote-fetch capability increases the risk of SSRF-like access patterns, data exfiltration, and policy bypass if callers assume the tool is strictly local.
The documentation frames vl_caller.py around API configuration, API failures, and raw provider output, which contradicts the manifest's presentation of a local PaddleOCR parsing skill. This is not just omitted detail: the file actively documents a provider-response wrapper and API error model inconsistent with a purely local parser.
The dependency specification uses a lower-bound version constraint instead of pinning to a specific version, which makes builds non-reproducible and can unexpectedly pull in vulnerable or breaking releases over time. In a document parsing skill that processes untrusted PDFs and images, dependency drift increases supply-chain and parser-exposure risk because image libraries are part of the attack surface.
# Install with: pip install -r scripts/paddleocr-doc-parsing/requirements-optimize.txt
# Image processing
Pillow>=10.0.0
# PDF processing
pypdfium2>=4.0.0
Pillow has multiple known advisories, and because the manifest does not pin a specific version, there is no way to verify whether deployed environments will install a safe or vulnerable release. This is more concerning in this skill's context because it processes document images, so Pillow likely handles untrusted input directly and could be exposed to malformed files that trigger memory corruption, resource exhaustion, or code execution vulnerabilities.
The pypdfium2 dependency is also unpinned, so installations may resolve to different versions across environments or over time. Because this skill parses PDFs locally, the PDF rendering/parsing library is security-sensitive, and version drift can expose the application to parser bugs or newly introduced issues.
Pillow>=10.0.0
# PDF processing
pypdfium2>=4.0.0
Using an unpinned dependency (httpx>=0.24.0) makes builds non-reproducible and can cause the environment to install unexpectedly vulnerable or breaking future versions. In a security-sensitive agent skill, this increases supply-chain risk because the exact package version cannot be audited or reliably constrained.
# PaddleOCR Document Parsing Dependencies
# HTTP client
httpx>=0.24.0
The manifest includes httpx without pinning to a version, while known advisories exist for some httpx releases. Because the installed version is not fixed, the deployment may resolve to an affected release, creating avoidable exposure to known dependency flaws.
The tool silently writes parsed document results to disk by default in the system temp directory, which may retain sensitive OCR output beyond the user's expectation. For document-processing skills, this can expose extracted text and structure to other local users, backup systems, or later processes if temp storage is not adequately protected or cleaned up.
No suspicious patterns detected.