Back to skill

Security audit

pdf-miner

Security checks for vulnerabilities and agentic risk

Overview

This PDF skill does useful extraction work, but it needs review because ordinary extraction can automatically send low-text PDF pages to an external vision API despite a top-level claim that OCR is out of scope.

Review before installing. Use this only if you are comfortable with low-text PDF pages being rendered and sent to a configured vision API when OCR is available, or run with --no-auto-ocr and avoid storing API keys in config.json. Do not use remote OCR on confidential PDFs unless that external processing is approved.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/extract_pdf.py:514
Finding

Automatic External Transmission of PDF Page Content Without Invocation-Level Consent

Content
View full analysis
env var > skill config.json > hardcoded default DEFAULT_OCR_API_KEY = os.environ.get("OCR_API_KEY", os.environ.get("OPENROUTER_API_KEY", _skill_cfg[0])) DEFAULT_OCR_BASE_URL = os.environ.get("OCR_BASE_URL", _skill_cfg[1] or "https://openrouter.ai/api/v1") DEFAULT_OCR_MODEL = os.environ.get("OCR_MODEL", _skill_cfg[2] or "qwen/qwen3.6-plus:free") ``` ```python def ocr_pdf_pages(pdf_path: str, page_nums: list, dpi: int = 200, api_key: str = None, base_url: str = None, model: str = None) -> dict[int, str]: """OCR specific pages of a PDF using the configured vision model.""" if not HAS_OCR: print(" Warning: OCR dependencies missing. Install: python -m pip install pymupdf openai") return {} api_key = api_key or DEFAULT_OCR_API_KEY base_url = base_url or DEFAULT_OCR_BASE_URL model = model or DEFAULT_OCR_MODEL # Auto-detect vision support and fallback if needed if not _is_vision_model(model): print(f" Warning: Model '{model}' may not support vision; falling back to default vision model.") model = "qwen/qwen3.6-plus:free" if not api_key: print(" Error: No API key. OCR skipped. Set OCR_API_KEY or use --ocr-api-key.") return {} client = OpenAI(api_key=api_key, base_url=base_url) doc = fitz.open(pdf_path) ocr_prompt = ( "Extract ALL visible text from this PDF page image. " "Preserve layout, tables, numbers exactly. Return ONLY the text." ) results = {} for pg_num in page_nums: try: page = doc[pg_num - 1] zoom = dpi / 72 pix = page.get_pixmap(matrix=fitz.Matrix( ...[truncated 5240 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Note
Location
config.json:1
Finding

Persistent Plaintext API Credentials in Project-Local Configuration

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
SKILL.md:19
Finding

Unpinned Third-Party Dependency Installation Instructions

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (16)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

This finding highlights a broader trust problem: the stated purpose emphasizes native PDF extraction, but the documentation introduces external OCR behavior and may overstate implemented features. Security reviewers and users rely on the manifest to understand capability boundaries; inaccurate claims can hide data egress and unexpected processing paths.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

This finding highlights a broader trust problem: the stated purpose emphasizes native PDF extraction, but the documentation introduces external OCR behavior and may overstate implemented features. Security reviewers and users rely on the manifest to understand capability boundaries; inaccurate claims can hide data egress and unexpected processing paths.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The file explicitly says the skill is not for OCR on scanned/image PDFs, but later documents built-in OCR support. Such internal contradiction can cause unsafe operator assumptions, especially for sensitive documents where OCR implies rendering pages and potentially sending them to third-party services.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The documentation says OCR is out of scope, then later states automatic OCR is enabled by default for low-text pages. Default activation of an undeclared remote-capable feature is particularly risky because users may trigger it unintentionally merely by processing a scanned or low-text PDF.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill explicitly implements OCR for scanned/image-based PDFs, which contradicts the stated skill boundary that it is 'NOT for OCR on scanned image-based PDFs.' This mismatch is dangerous because it expands capability beyond what users and reviewers are led to expect, and in this implementation OCR is not purely local but tied to external API transmission of page images.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The OCR path renders PDF pages to JPEG, base64-encodes them, and sends them to an external vision API via chat completions. For a PDF extraction skill that users would reasonably expect to run locally, this creates a significant confidentiality risk: sensitive report pages, financial documents, contracts, or research PDFs may be exfiltrated to a third party without prominent disclosure.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The manifest explicitly says the skill is 'NOT for: OCR on scanned image-based PDFs', but this file is dedicated to OCR of scanned/image-based PDF pages using page-to-image conversion and a vision API. That is a direct description-behavior mismatch at the skill level, not merely an implementation detail.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
81% confidence
Finding

The skill documentation instructs use of environment variables and local file reads/writes, but the manifest does not declare any tool scope or permissions. That gap weakens reviewability and least-privilege controls, making it easier for a caller to overlook that the skill accesses credentials and writes extracted output files.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill introduces external vision-API use with API credentials for a task presented as simple PDF extraction. That expands the trust boundary to a third party and creates confidentiality risk for uploaded page images/text, especially with financial or research PDFs that may contain sensitive data.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The OCR workflow sends PDF page content to an external vision API, but the documentation lacks a prominent privacy and data-transmission warning. In this skill context, PDFs may contain proprietary financial, research, or customer data, so silent third-party transmission materially increases confidentiality and compliance risk.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The documentation describes OCR as optional or flag-driven, but the implementation enables automatic OCR by default through --auto-ocr defaulting to true and do_ocr = args.ocr or args.auto_ocr. This is dangerous because users may invoke ordinary PDF extraction and unknowingly trigger external processing of low-text pages, violating expectations and potentially exposing sensitive data.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The code reads OCR-related secrets from environment variables and local config to enable outbound service access, which is outside the core expectation of simple PDF parsing. While reading secrets is not inherently malicious, here it supports undeclared networked processing of document contents and makes silent external integration easier.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The code sends page images to a third-party OCR endpoint without a clear user-facing warning at the moment of transmission. In the context of a document-mining skill, PDFs commonly contain confidential business, legal, or personal data, so lack of explicit disclosure materially increases privacy and compliance risk.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The code transmits rendered PDF page images to a configurable external API endpoint via an OpenAI-compatible client. PDFs often contain sensitive business, financial, legal, or personal data, so sending full page images off-host can cause confidentiality breaches, especially because the base URL is configurable and may point to untrusted infrastructure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The OCR workflow sends page image data to an external vision API without any explicit warning, consent gate, or privacy notice in the code path. In the context of a PDF extraction skill, users may reasonably expect local processing, so silent remote transmission increases the risk of unintended disclosure of sensitive document contents.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
81% confidence
Finding

The manifest emphasizes reading/extracting content from PDFs, but this implementation persists results to a new markdown file by default. Creating output files is a broader operational behavior than a pure extractor/reader description and is not called out in the manifest.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.