Back to skill

Security audit

pdf

Security checks for vulnerabilities and agentic risk

Overview

This PDF skill is a coherent local PDF-processing guide with helper scripts, but users should be careful with password-protected documents, overwrite-prone commands, and unpinned package installs.

Install or use this skill only for PDFs you are authorized to process. Prefer a virtual environment with pinned dependencies, avoid putting real passwords directly in shell commands when possible, and write repaired or filled PDFs to new output files so originals are not overwritten.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:215
Finding

Unpinned Third-Party Dependency Installation

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, line 215
Vulnerability Type: T08: Insecure Dependencies
Risk Level: Medium

Complete Code Snippet:

python
# Requires: pip install pytesseract pdf2image
import pytesseract
from pdf2image import convert_from_path

Technical Analysis

The documentation directs users to install third-party packages without version constraints, cryptographic hashes, a lock file, or an explicitly trusted package index. Consequently, package names may resolve to mutable releases and uncontrolled transitive dependencies from the user's configured Python package index.

Python package installation may execute package build or installation logic. A compromised future release, compromised transitive dependency, or package substitution through an untrusted index could therefore cause attacker-controlled code to run. The project also imports additional third-party libraries but does not provide a dependency manifest that records reviewed versions and artifact hashes.

No evidence indicates that pytesseract or pdf2image is currently malicious. This finding concerns the unsafe and non-reproducible installation method rather than confirmed malicious package content.

Attack Path

  1. A user follows the instruction in SKILL.md and runs pip install pytesseract pdf2image.
  2. pip queries the package index configured in the user's environment.
  3. The resolver selects current package releases and transitive dependencies without validating them against project-approved versions or hashes.
  4. A compromised release, malicious transitive dependency, or substituted distribution is downloaded.
  5. Attacker-controlled code executes during installation or when the dependency is subsequently imported by the PDF-processing workflow.

Impact Assessment

Malicious dependency code would execute with the privileges of the user installing or running the Skill. It could potentiall ...[truncated 464 chars]

Remediation
View remediation

Remediation Suggestions

  1. Add a reviewed dependency manifest with exact versions for all third-party Python dependencies.

  2. Generate and record cryptographic hashes for every direct and transitive distribution.

  3. Require hash verification during installation, for example:

    bash
    python -m pip install --require-hashes -r requirements.txt
    
  4. Use a lock-generation process that resolves transitive dependencies reproducibly.

  5. Document the trusted package index and prevent fallback to uncontrolled extra indexes.

  6. Prefer prebuilt, reviewed artifacts and avoid source builds where practical.

  7. Automate dependency vulnerability scanning and periodic review before updating pinned versions.

  8. Perform installation in an isolated virtual environment under a non-privileged account.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (7)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The supplied code does not implement or expose the declared PDF toolkit capabilities such as extracting text or tables, creating PDFs, merging/splitting files, or filling/analyzing PDFs directly. Instead, it is a test script for a helper function that validates form field bounding boxes and related layout constraints using synthetic JSON inputs. While this may be tangentially related to PDF form handling, the chunk’s actual purpose is much narrower and materially different from the declared comprehensive PDF manipulation description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description presents a toolkit centered on PDF manipulation tasks such as extracting text and tables, creating PDFs, merging/splitting documents, and handling forms. The supplied code instead performs PDF page rasterization to PNG files, which is a distinct capability not mentioned in the description. While PDF-to-image conversion is related to PDF processing broadly, it is not reasonably covered by the specific declared functions listed, and it is the sole primary behavior of this code chunk. Therefore the code's actual behavior is materially different from the declared purpose.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The guide provides a concrete command for decrypting password-protected PDFs without any authorization, ownership, or legal-use warning. In an agent skill, this can normalize or facilitate bypassing document protections on third-party files, increasing the risk of unauthorized access to sensitive information.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This markdown file includes commands for adding and removing PDF encryption, including an inline password example, but does not warn users that these actions involve sensitive credentials and access-controlled content. Under the markdown-specific warning criterion, behaviors affecting privacy or protected data should be accompanied by a clear caution.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The troubleshooting section recommends qpdf --replace-input corrupted.pdf, which overwrites the original file, but the markdown does not warn that this operation is destructive to the source document. For markdown files, user-facing descriptions should disclose behaviors that can affect user data or system integrity.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

This code performs a file write operation to an arbitrary output path supplied on the command line. While the script prints a success message afterward, it does not warn beforehand that an existing file may be overwritten, and the surrounding comments do not disclose that risk.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.