Back to skill

Security audit

PaddleOCR Text Recognition

Security checks for vulnerabilities and agentic risk

Overview

This OCR skill appears purpose-built for text extraction, but users should review it because local documents are sent to an external OCR API without a clear privacy confirmation.

Install only if you are comfortable using PaddleOCR's external API for the images or PDFs you submit. Avoid sending secrets, regulated records, internal screenshots, or confidential documents unless that service is approved for that data, and prefer a pinned dependency or controlled runtime for higher-risk use.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:20
Finding

Unpinned Third-Party Executable Dependency

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 20-23
Vulnerability Type: T08: Insecure Dependencies
Risk Level: Medium

yaml
install:
  - kind: uv
    package: paddleocr
    bins: [paddleocr]

Technical Analysis

The skill directs the runtime to install the paddleocr package without specifying an exact version or an integrity hash. Consequently, installation may resolve to a future package release whose contents differ from the version reviewed during this audit.

Because the installed package provides an executable that the skill invokes, malicious package initialization or runtime code could execute with the permissions of the agent process. This creates a supply-chain risk if the package repository, publisher account, release process, or dependency resolution path is compromised.

Attack Path

  1. An attacker compromises the package publisher, distribution repository, or a future dependency release.
  2. The attacker publishes malicious code under a version permitted by the unpinned package specification.
  3. The skill installation process resolves and installs that release.
  4. Malicious installation or runtime code executes when the package is installed or the paddleocr executable is invoked.
  5. The code can access resources available to the agent process, potentially including submitted documents and the PADDLEOCR_ACCESS_TOKEN.

Impact Assessment

Successful exploitation could permit arbitrary code execution under the agent process's existing privileges. The resulting scope could include reading files accessible to that process, accessing environment variables such as the OCR access token, modifying writable project data, and making outbound network requests. The configuration does not itself establish elevated privileges or persistence, so impact remains bounded by the permissions and isolation controls of the runtime environment.

Remediation
View remediation

Remediation Suggestions

  • Pin paddleocr to an exact, reviewed version rather than allowing unconstrained resolution.
  • Enforce package integrity verification using trusted hashes, signatures, or a lock file supported by the installation environment.
  • Retrieve packages only from an explicitly configured and trusted package index.
  • Review both direct and transitive dependencies before updating the pinned version.
  • Run the CLI in a sandbox with minimal filesystem access, restricted environment-variable exposure, and constrained outbound network access.
  • Separate the API token from package installation processes wherever the runtime permits it.

other

Warning
Location
SKILL.md:52
Finding

Local Documents Transmitted to an External OCR Service Without Explicit Disclosure

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 52-58
Vulnerability Type: other: Unwarned External Data Transmission
Risk Level: Medium

bash
From local file:

```bash
paddleocr api \
  --model_type ocr \
  --file_path "./document.pdf"
text

### Technical Analysis

The documented command submits a local file through the `paddleocr api` interface for hosted OCR processing. Although the skill requires an external-service access token, it does not clearly disclose at the point of use that local images or PDFs will leave the local environment. It also provides no instruction to obtain confirmation before transmitting potentially confidential material.

Files processed by this workflow may contain personal information, credentials, financial records, internal documents, or other sensitive content. Without an explicit disclosure and approval boundary, a user may reasonably expect local-file OCR to occur locally and may not understand the external service's retention, logging, jurisdiction, or secondary-processing implications.

### Attack Path

1. A user requests OCR of a sensitive local image or PDF.
2. The agent follows the documented `--file_path` workflow.
3. The local document is submitted to the external OCR service.
4. Document contents become available to systems and operators governed by the external provider's security, retention, and access policies.
5. If the provider, transport path, service account, or returned result location is compromised or improperly configured, sensitive document content may be exposed.

### Impact Assessment

The primary impact is loss of confidentiality for the complete contents of submitted documents. Exposure scope is limited to files deliberately supplied to the OCR command, but those files may contain highly sensitive information. This behavior does not by itself grant local privilege escalation, persistence, or unrestricted filesystem access.
Remediation
View remediation

Remediation Suggestions

  • State explicitly that the paddleocr api command uploads local document contents to an external OCR provider.
  • Require informed user confirmation before transmitting each local file or batch of files.
  • Identify the service destination and link to applicable privacy, retention, deletion, and data-processing policies.
  • Warn users not to submit secrets or regulated data unless the service has been approved for that data classification.
  • Provide a local-only OCR alternative for sensitive documents.
  • Minimize submitted data by limiting pages or cropping images when full-document processing is unnecessary.
  • Avoid logging document contents, tokens, or complete service responses unless explicitly required.
  • Ensure transport encryption is enforced and restrict returned result URLs from being exposed beyond the requesting session.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (3)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest trigger list includes broad terms such as "screenshot," "photo scan," and especially "recognize text," which can appear in ordinary conversation outside the intended OCR-skill context. Although some terms are domain-specific, the list does not define boundaries or negative examples to clarify when this skill should or should not activate.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill instructs users to send local files or remote URLs to an external OCR API but does not warn that document contents may leave the local environment and be processed by a third party. This can lead to unintended disclosure of sensitive documents, screenshots, or internal URLs, especially when users assume OCR is happening locally.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The instruction to 'always show the full extracted content' encourages verbatim disclosure of all OCR'd text, which may include passwords, IDs, financial data, medical information, or other secrets present in screenshots and scans. In an OCR skill, this is especially risky because users often submit raw document images that contain more sensitive text than they realize.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.