Back to skill

Security audit

pdf

Security checks for vulnerabilities and agentic risk

Overview

This PDF skill appears to be a normal PDF-processing helper, with a few cautionary examples around OCR installs, decryption, and in-place repair.

Install only if you are comfortable letting the skill process PDFs you choose. Use copies for repair operations, decrypt only PDFs you are authorized to access, and install OCR dependencies in a dedicated environment with reviewed or pinned versions when possible.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:215
Finding

Unpinned Third-Party Python Package Installation

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, line 215
Vulnerability Type: Unpinned third-party dependencies
Risk Level: Medium

Vulnerable Code

python
# Requires: pip install pytesseract pdf2image

Technical Analysis

The OCR workflow directs users or agents to install pytesseract and pdf2image directly from the package index without specifying reviewed versions, cryptographic hashes, a lockfile, or an authenticated dependency manifest.

Because package resolution depends on mutable registry state, future installations may retrieve releases or transitive dependencies that differ from those originally reviewed. If a package publisher account, package release, or dependency is compromised, malicious code could execute during installation or when the dependency is imported and used.

The audit found no evidence that either named package is currently malicious. The finding concerns the unsafe, non-reproducible installation mechanism.

Attack Path

  1. A user or agent follows the OCR instructions in SKILL.md.
  2. It runs pip install pytesseract pdf2image.
  3. pip resolves the latest available package versions and their transitive dependencies.
  4. If any resolved artifact has been compromised, its installation behavior or imported code executes locally.
  5. The malicious component operates with the permissions and data access of the Python environment running the Skill.

This path requires compromise of a named package, its publisher account, a transitive dependency, or the configured package-index infrastructure.

Impact Assessment

Successful supply-chain exploitation could provide arbitrary code execution with the privileges of the user running pip or the PDF workflow. Accessible scope could include local PDF contents, generated outputs, files readable by that user, environment variables, and network resources available to the process.

No privilege-escalation mechanism, persistence mechanism, credential harvestin ...[truncated 71 chars]

Remediation
View remediation

Remediation Suggestions

  1. Pin every direct dependency to a reviewed, exact version.
  2. Commit a dependency lockfile that also fixes transitive dependency versions.
  3. Require cryptographic hashes for all downloaded artifacts, such as through a hashed requirements file and pip install --require-hashes.
  4. Configure installation to use an explicitly approved package index.
  5. Install dependencies inside a dedicated, least-privileged virtual environment or sandbox.
  6. Regularly scan and update pinned dependencies after security review.
  7. Document required non-Python components, including Tesseract OCR and Poppler, with trusted installation sources and reviewed versions.

A hardened example would use a reviewed requirements file:

bash
python -m pip install --require-hashes -r requirements.txt
Vulnerability Patterns
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (5)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The supplied code chunk is a test module (check_bounding_boxes_test.py) that exercises get_bounding_box_messages using synthetic form field data. It checks for bounding box intersections, page separation, font-size/height validation, and output limiting. This is materially narrower and different from the declared description of a comprehensive PDF manipulation toolkit. While bounding-box/form validation could be a supporting part of PDF form handling, this chunk itself does not demonstrate the broad declared capabilities such as extracting text/tables, creating PDFs, merging/splitting documents, or general PDF processing. Therefore the description does not accurately represent what this specific code chunk actually does.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

The qpdf example explicitly shows decrypting a password-protected PDF but provides no warning about authorization, legal constraints, or sensitive-data handling. In an agent skill, this can normalize removal of document protections and lead automated workflows to process confidential content without appropriate user confirmation or policy checks.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The troubleshooting section recommends qpdf --replace-input corrupted.pdf, which overwrites the original file in place. The markdown does not warn that this operation modifies the only copy of the user's PDF and may risk data loss if the repair attempt fails or produces unwanted results.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

The manifest describes extracting text and tables, creating and modifying PDFs, and handling forms, all centered on PDF-native processing. This section introduces OCR via pytesseract and pdf2image, which is a separate image-recognition capability rather than an obvious PDF manipulation primitive.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.