Back to skill

Security audit

OCR with python

Security checks for vulnerabilities and agentic risk

Overview

The OCR skill appears purpose-aligned, but it uses unpinned package installs and insecure shared temporary files when handling potentially sensitive documents.

Review before installing. Use a pinned, reviewed requirements file or controlled package source, and avoid processing sensitive PDFs on shared machines until the temporary-file handling is fixed to use a private per-run directory. Expect Chinese OCR behavior unless the language setting is made configurable.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:21
Finding

Unpinned Third-Party Dependencies Permit Unreviewed Package Changes

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:21-25 and SKILL.zh-CN.md:21-25
Vulnerability Type: Unpinned package installation
Risk Level: Medium

The installation instructions contain the following command:

bash
pip3 install paddlepaddle paddleocr

Technical Analysis

The dependencies are installed without version constraints, integrity hashes, or a lock file. Consequently, each installation resolves whatever versions the configured Python package index currently considers appropriate.

Although the referenced package names appear consistent with the documented OCR functionality, the absence of version and integrity controls means that the installed code can change after the Skill has been reviewed. Python packages may execute code during installation and are imported by scripts/ocr.py, so a compromised, malicious, or unexpectedly incompatible future release would execute with the permissions of the user running the installation or OCR script.

The equivalent instruction is present in both the English and Chinese documentation.

Attack Path

  1. An attacker compromises a dependency release or the package-distribution account used to publish it.
  2. The attacker publishes a malicious version under one of the dependency names.
  3. A user follows the documented unpinned installation command.
  4. pip resolves and installs the attacker-controlled release.
  5. Malicious code executes during package installation or when paddleocr is imported by the OCR script.

This path depends on compromise of the selected package source or release process; the audited repository does not itself host an embedded malicious dependency.

Impact Assessment

Malicious dependency code would run with the privileges of the user executing pip or the OCR script. It could access files, credentials, environment variables, and network resources available to that user, alter user-owned data, or execute additional pr ...[truncated 132 chars]

Remediation
View remediation

Remediation Suggestions

  • Pin every direct dependency to a reviewed version, for example through a version-controlled requirements file.
  • Pin transitive dependencies with a lock-generation tool appropriate for Python.
  • Require package hashes by installing with pip install --require-hashes -r requirements.txt.
  • Use a trusted, explicitly configured package index and disable unexpected supplemental indexes.
  • Add automated dependency vulnerability and provenance scanning.
  • Review and update pinned versions through a controlled update process rather than resolving the latest versions during installation.
  • Update both SKILL.md and SKILL.zh-CN.md so that their installation instructions remain consistent.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/ocr.py:29
Finding

Predictable Shared Temporary Files Allow Symlink and Collision Attacks

Content
View full analysis

Vulnerability Details

File Location: scripts/ocr.py:29-31
Vulnerability Type: Insecure temporary-file creation
Risk Level: Medium

The PDF extraction function writes images to deterministic names in the shared /tmp directory:

python
output_path = f"/tmp/pdf_page{page_num+1}_img{img_index}.{image_ext}"
with open(output_path, "wb") as f:
    f.write(image_bytes)
images.append(output_path)

Temporary-file cleanup also suppresses all errors:

python
for img_path in images:
    try:
        os.remove(img_path)
    except:
        pass

Technical Analysis

Temporary filenames are based only on page number, image index, and image extension. They contain no process-specific or cryptographically random component. Multiple executions processing common image formats therefore use the same paths.

open(output_path, "wb") follows symbolic links and does not request exclusive file creation. On a multi-user system where another local user can create entries in /tmp, an attacker can pre-create a predicted path as a symbolic link. When the victim runs the script, Python follows that link and opens the target with truncation enabled. The target must still be writable by the victim process, but the attacker may redirect the extracted bytes to any such target.

Predictable names also cause concurrent OCR executions to overwrite, read, or delete one another's temporary images. Because cleanup catches every exception without reporting it, failures can leave extracted document images on disk without notifying the user.

Attack Path

  1. A local attacker observes or predicts that the victim will process a PDF containing a common embedded image format such as PNG or JPEG.
  2. The attacker creates a symbolic link at a predictable path, such as /tmp/pdf_page1_img0.png, pointing to a file writable by the victim.
  3. The victim invokes scripts/ocr.py on the PDF.
  4. The script opens the predi ...[truncated 1139 chars]
Remediation
View remediation

Remediation Suggestions

  • Create a private temporary directory with tempfile.TemporaryDirectory() for each invocation.
  • Store all extracted images beneath that randomly named directory rather than directly under /tmp.
  • Use context managers so cleanup occurs on success and exceptions.
  • If individual temporary files are necessary, use tempfile.NamedTemporaryFile() or secure exclusive creation rather than constructing filenames manually.
  • Do not follow attacker-controlled symbolic links; secure temporary-file APIs provide the preferred protection.
  • Ensure cleanup failures are logged rather than suppressed with a bare except.
  • Avoid exposing complete temporary paths in routine output if those paths may assist local race attempts.
  • Add concurrency tests verifying that simultaneous OCR jobs cannot overwrite or delete each other's files.
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (3)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The example initializes PaddleOCR with lang='ch', which forces Chinese mode in the documented usage. Under the policy, locale or language should not be fixed without user opt-in unless clearly justified; this file says the skill supports both Chinese and English, but the sample only enforces Chinese and does not explain how to choose English.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The code initializes PaddleOCR with lang='ch', forcing a specific language/locale for all users. This is a natural-language policy concern because the skill does not offer opt-in, configuration, or documentation explaining why only Chinese OCR is used.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The natural-language description states that the skill supports Chinese and English text extraction, and the body repeats that it supports both Chinese and English, which constrains language handling without indicating user opt-in or explaining why other languages are excluded. Under the language/locale policy, fixed language constraints should either offer a choice or be clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.