Back to skill

Security audit

pdf

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent PDF toolkit, but it needs review because some instructions can expose sensitive form data or overwrite PDFs without enough safeguards.

Review before installing. Use this skill only on PDFs you are authorized to access, keep form JSON/images/filled PDFs in approved local locations, delete sensitive intermediates when done, avoid copying the in-place qpdf --replace-input repair command unless you have a backup, and install optional OCR dependencies in an isolated environment with trusted, pinned packages where possible.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:215
Finding

Unpinned Third-Party Dependency Installation Instruction

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 215–218
Vulnerability Type: Unpinned dependencies resolved from a mutable package registry
Risk Level: Medium

Vulnerable Code Snippet:

python
# Requires: pip install pytesseract pdf2image
import pytesseract
from pdf2image import convert_from_path

Technical Analysis

The OCR example instructs users or agents to install pytesseract and pdf2image without exact version constraints, package hashes, a lockfile, or an explicitly trusted package index. Consequently, the installed code is determined by the state and configuration of the package registry at installation time rather than by the audited project contents.

Although the referenced package names are not inherently malicious, an unpinned installation may retrieve newly released, compromised, or otherwise incompatible package versions and their transitive dependencies. Python packages can execute code during installation and subsequently when imported. This creates a supply-chain trust boundary that is not controlled by the project.

Attack Path

  1. A user requests OCR processing of a scanned PDF.
  2. A user or agent follows the instruction in SKILL.md and runs pip install pytesseract pdf2image.
  3. pip resolves the latest acceptable packages and transitive dependencies from the environment's configured index.
  4. If the index, a package release, or a transitive dependency has been compromised, attacker-controlled code is installed.
  5. Malicious code may execute during installation or when the documented imports are evaluated.
  6. The code runs with the permissions of the account performing the installation or invoking the OCR workflow.

Impact Assessment

Successful exploitation could permit arbitrary code execution with the invoking user's privileges. Depending on that account's permissions and environment, the affected scope could include accessible project files, proc ...[truncated 386 chars]

Remediation
View remediation

Remediation Suggestions

  1. Replace the inline unpinned installation command with a reviewed dependency manifest containing exact versions.
  2. Generate and verify cryptographic hashes for every direct and transitive dependency, then require hash validation during installation, for example with pip install --require-hashes -r requirements.txt.
  3. Use a lockfile generated from a controlled build process and update it through reviewed dependency-upgrade changes.
  4. Explicitly configure an approved package index rather than relying on arbitrary environment-level index settings.
  5. Run dependency installation and PDF processing in an isolated virtual environment or restricted container using a non-privileged account.
  6. Add automated dependency vulnerability and integrity scanning to the maintenance workflow.
  7. Document the required external OCR executable separately, including trusted installation sources and supported versions.
Vulnerability Patterns
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (6)

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The supplied code does not demonstrate a comprehensive PDF manipulation toolkit. Instead, it is a unit test module for a specific helper, get_bounding_box_messages, that validates spatial relationships among form-field bounding boxes and related font-size constraints. While this is adjacent to PDF form handling, the actual code chunk is narrowly about QA/validation logic over JSON-formatted field metadata, not extracting text/tables, generating PDFs, merging/splitting documents, or broad PDF processing. That makes the observed behavior materially narrower and different from the declared primary purpose.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill explicitly documents PDF password removal and decryption without any caution about authorization, sensitive data handling, or legal and policy constraints. In an agent context, this can normalize or enable processing of protected documents without confirming the user is permitted to decrypt them, increasing the risk of unauthorized access to confidential information.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The workflow instructs users to generate field_values.json and a filled output PDF that are likely to contain sensitive personal data, but it provides no warning about local persistence, overwrite risk, or downstream handling of those artifacts. In an agent context, this can lead to unintended storage of PII in working directories, logs, synced folders, or reused temp locations.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The non-fillable workflow creates multiple artifacts—fields.json, validation images, and annotated output PDFs—that can directly expose sensitive form contents and visually reveal personal information. Because the instructions mandate iterative creation and inspection of these files without privacy or retention safeguards, the chance of accidental disclosure is elevated.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The reference recommends qpdf --replace-input corrupted.pdf, which overwrites the original file in place. In a skill intended for automated PDF processing, an agent or user may copy this command directly and irreversibly modify the only available copy of a damaged document, causing data loss or preventing later forensic/recovery attempts. The PDF-processing context makes this more dangerous because users are likely to apply repair commands to important source documents at scale.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.