Back to skill

Security audit

pdf

Security checks for vulnerabilities and agentic risk

Overview

This skill is a disclosed PDF processing guide with local helper scripts, with some documentation cautions missing but no evidence of hidden or malicious behavior.

Install and use this skill for PDF work only in a normal project or scratch directory. Keep backups before using examples that modify PDFs, avoid hardcoding real passwords in commands, and install optional Python dependencies in a virtual environment with reviewed versions when possible.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:215
Finding

Unpinned Third-Party Python Dependencies

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:215-217
Vulnerability Type: T08: Insecure Dependencies
Risk Level: Medium

Vulnerable Code Snippet:

python
# Requires: pip install pytesseract pdf2image
import pytesseract
from pdf2image import convert_from_path

Technical Analysis

The installation guidance uses pip install pytesseract pdf2image without pinned versions, cryptographic hashes, a dependency lockfile, or an explicitly trusted package index. Package resolution therefore depends on mutable external package-index state at installation time.

The named packages are established packages rather than apparent typosquatting attempts, and the project does not automatically execute this installation command. Nevertheless, an agent following the documented workflow could install versions that differ from those reviewed by the project author. If a package release, transitive dependency, configured package index, or package-resolution environment is compromised, installation or subsequent import may execute attacker-controlled code.

Attack Path

  1. A user requests OCR processing of a scanned PDF.
  2. The agent follows the installation instruction in SKILL.md.
  3. pip resolves the latest compatible versions from its configured indexes because no versions or hashes are specified.
  4. An attacker compromises a resolved package or transitive dependency, or controls a higher-priority configured package index.
  5. The malicious component executes during installation or when the package is imported by the OCR workflow.
  6. The component gains access to the files, environment variables, network access, and operating-system permissions available to the account running the agent.

Impact Assessment

Successful exploitation could provide arbitrary code execution with the privileges of the user or service account running pip and the PDF workflow. The resulting scope may include access to PD ...[truncated 390 chars]

Remediation
View remediation

Remediation Suggestions

  1. Pin every direct dependency to a reviewed version, for example:
    bash
    python -m pip install --index-url https://pypi.org/simple pytesseract==REVIEWED_VERSION pdf2image==REVIEWED_VERSION
    
  2. Commit a requirements or lock file that also constrains transitive dependencies.
  3. Generate and verify cryptographic hashes, then install with --require-hashes.
  4. Use an approved internal package mirror or explicitly specify the trusted official index to reduce dependency-confusion exposure.
  5. Install packages inside a dedicated virtual environment or sandbox under an unprivileged account.
  6. Disable unnecessary network and filesystem access while processing untrusted PDFs.
  7. Add automated dependency vulnerability and integrity scanning to the release process.
  8. Document tested dependency versions and periodically update them through a controlled review process.
Vulnerability Patterns
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (8)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents a broad PDF manipulation toolkit. The supplied code chunk, however, is a test file for get_bounding_box_messages, focused on validating layout rules for form-field bounding boxes in JSON-like input. This is related at most tangentially to PDF forms, but it does not implement the declared primary capabilities such as extracting text or tables, creating PDFs, merging/splitting documents, or filling forms. Because the code’s actual purpose is a narrow test harness for bounding-box validation rather than a comprehensive PDF toolkit, this chunk does not accurately represent the declared description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
91% confidence
Finding

The declared description broadly covers several PDF manipulation tasks such as extraction, creation, merging/splitting, and forms handling, but this specific code performs rasterization of PDF pages into PNG files. That is a distinct capability not explicitly represented in the description. While it is adjacent to PDF processing, the primary behavior here is image export rather than the listed text/table extraction, PDF creation, document splitting/merging, or form handling.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill demonstrates file-writing capability by creating and modifying PDFs, exporting Excel files, and writing text outputs, but it does not declare any tool scope such as permissions or allowed-tools. In an agent environment, undeclared write capabilities weaken policy enforcement and can enable unintended filesystem modification if the skill is invoked on attacker-controlled inputs or in sensitive directories.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The reference includes commands for adding and removing PDF password protection, including an inline password example, but provides no warning about authorized use, sensitive password handling, or the implications of decrypting protected files. In markdown files, safety-affecting behaviors that may impact privacy or protected data should be disclosed to the user.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The qpdf --replace-input corrupted.pdf example performs an in-place write to the source file, which can affect user data and is not accompanied by any caution or backup recommendation. For markdown skill documentation, destructive or integrity-affecting behaviors should be clearly disclosed.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The manifest describes capabilities for extracting text and tables, creating PDFs, merging/splitting documents, and handling forms, but this script converts PDF pages into PNG image files. Rendering PDFs to images is a distinct document-to-image conversion capability that is not reflected in the stated description.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

This markdown file directs the user to run scripts that generate output PDFs and intermediate JSON files, which are file-writing operations. While the commands are explicit, the description does not include any caution that these steps will create or overwrite local files or advise choosing safe output paths.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

These instructions tell the user to generate validation images and a filled PDF, which affects local data and produces new artifacts. The skill description does not warn that these operations will write files or recommend preserving the original PDF before annotation.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.