Back to skill

Security audit

PR's PDF Agent

Security checks for vulnerabilities and agentic risk

Overview

This PDF tool mostly matches its stated purpose, but it needs review because some features can fetch remote URLs, send document text to LLM commands, and expose PDF passwords in process details or errors.

Review before installing in environments with sensitive PDFs, untrusted URLs, or shared hosts. Prefer local file inputs, avoid using --llm-cmd or translation on confidential documents unless you trust the command/provider, do not pass reusable passwords, and avoid rasterized redaction unless intermediate files are cleaned securely.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
pdfagent/tools/html_to_pdf.py:90
Finding

Unrestricted Remote HTML Fetching Enables Server-Side Request Forgery

Content
View full analysis
str: source_path = Path(source) if source_path.exists(): return source_path.read_text(encoding="utf-8") if source.startswith("http://") or source.startswith("https://"): with urlopen(source) as resp: return resp.read().decode("utf-8", errors="ignore") return source ``` The vulnerable functionality is directly exposed through the CLI: ```python @app.command("html-to-pdf") def html_to_pdf_cmd( source: str = typer.Argument(..., help="URL or HTML file path"), out: Path = typer.Option(..., "--out"), json_output: bool = typer.Option(False, "--json"), usage_file: Optional[Path] = typer.Option(None, "--usage-file"), ): meter = UsageMeter("html_to_pdf", [Path(source)] if Path(source).exists() else []) try: html_to_pdf(source, out) ``` ### Technical Analysis The `source` argument is controlled by the caller and may contain an arbitrary HTTP or HTTPS URL. `_load_html()` passes that URL directly to `urllib.request.urlopen()` without validating the destination hostname or resolved IP address. There are no controls to prevent requests to: - Loopback addresses such as `127.0.0.1` or `::1`. - RFC 1918 private networks. - Link-local addresses. - Cloud instance metadata services. - Internal DNS names. - Reserved or otherwise non-public address ranges. Redirect destinations are not independently validated. Consequently, an initially public URL could redirect the request to a prohibited internal destination. The request also has no explicit connection or read timeout and no maximum response-size limit. The primary HTML conversion paths pass the same uncontrolled source to `wkhtmltopdf ...[truncated 1561 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
pdfagent/tools/security.py:9
Finding

PDF Passwords Are Exposed in Process Arguments and Error Messages

Content
View full analysis
None: qpdf = require_bin("qpdf", "Install qpdf to enable unlock") output_path.parent.mkdir(parents=True, exist_ok=True) cmd = [qpdf, f"--password={password}", "--decrypt", str(input_path), str(output_path)] run_cmd(cmd) ``` The protect operation similarly places both the user and owner passwords in process arguments: ```python def protect_pdf( input_path: Path, output_path: Path, user_password: str, owner_password: str | None = None, allow_print: bool = False, allow_copy: bool = False, allow_modify: bool = False, ) -> None: qpdf = require_bin("qpdf", "Install qpdf to enable protect") output_path.parent.mkdir(parents=True, exist_ok=True) owner = owner_password or user_password perms = [] if allow_print: perms.append("--print=full") else: perms.append("--print=none") if allow_copy: perms.append("--extract=y") else: perms.append("--extract=n") if allow_modify: perms.append("--modify=all") else: perms.append("--modify=none") cmd = [ qpdf, "--encrypt", user_password, owner, "256", *perms, "--", str(input_path), str(output_path), ] run_cmd(cmd) ``` The shared command runner includes the complete command in timeout and command-failure exceptions: ```python except subprocess.TimeoutExpired as exc: raise ExternalCommandError(f"Command timed out: {' '.join(cmd)}") from exc ...[truncated 2858 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (41)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

External LLM-based translation of PDF contents directly contradicts the self-hosted positioning and may transmit extracted document text to an external provider or service. In the context of a PDF-processing skill, this is especially dangerous because users are likely to feed it sensitive business, legal, or personal documents under the assumption that content remains local.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

External LLM-based translation of PDF contents directly contradicts the self-hosted positioning and may transmit extracted document text to an external provider or service. In the context of a PDF-processing skill, this is especially dangerous because users are likely to feed it sensitive business, legal, or personal documents under the assumption that content remains local.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

External LLM-based translation of PDF contents directly contradicts the self-hosted positioning and may transmit extracted document text to an external provider or service. In the context of a PDF-processing skill, this is especially dangerous because users are likely to feed it sensitive business, legal, or personal documents under the assumption that content remains local.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

External LLM-based translation of PDF contents directly contradicts the self-hosted positioning and may transmit extracted document text to an external provider or service. In the context of a PDF-processing skill, this is especially dangerous because users are likely to feed it sensitive business, legal, or personal documents under the assumption that content remains local.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

External LLM-based translation of PDF contents directly contradicts the self-hosted positioning and may transmit extracted document text to an external provider or service. In the context of a PDF-processing skill, this is especially dangerous because users are likely to feed it sensitive business, legal, or personal documents under the assumption that content remains local.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

External LLM-based translation of PDF contents directly contradicts the self-hosted positioning and may transmit extracted document text to an external provider or service. In the context of a PDF-processing skill, this is especially dangerous because users are likely to feed it sensitive business, legal, or personal documents under the assumption that content remains local.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

External LLM-based translation of PDF contents directly contradicts the self-hosted positioning and may transmit extracted document text to an external provider or service. In the context of a PDF-processing skill, this is especially dangerous because users are likely to feed it sensitive business, legal, or personal documents under the assumption that content remains local.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The _run_command_llm and _run_command_text paths execute an arbitrary external command from the cmd parameter, using shlex.split but no allowlisting, sandboxing, or policy restriction. Even though shell metacharacter injection is reduced by avoiding shell=True, this still permits execution of attacker-controlled or unsafe binaries and can be abused for arbitrary code execution, data access, or prompt exfiltration in a PDF-processing context where such power is unjustified.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill declares no explicit tool scope even though it appears capable of shell execution, filesystem access, environment access, and possible network use. That mismatch increases the chance an agent or operator will invoke it with broader privileges than expected, which can enable command execution, data exposure, or unintended outbound access.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The CLI explicitly accepts a URL or local HTML path for HTML-to-PDF conversion, which expands the tool from purely self-hosted local document processing into remote content fetching. That can create SSRF, unexpected network egress, or processing of untrusted remote content, especially in automated or server-side deployments where the operator may assume only local files are handled.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The agent command can invoke external LLM providers or arbitrary LLM commands, which materially broadens the trust boundary beyond local PDF tooling. In security-sensitive environments, this can cause unexpected network disclosure of prompts or execution of external helper commands not implied by a self-hosted PDF utility.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Natural-language agent mode sends user instructions through an LLM planning path when provider is not 'none', introducing remote/model-mediated behavior that is not core to PDF processing. This increases the risk of data disclosure, nondeterministic execution plans, and operator surprise in environments expecting deterministic local tooling.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The translate path can pass document-derived text to an external provider or command, yet this code path shows no user-facing warning before potentially exporting sensitive content. For PDFs containing confidential, regulated, or internal data, silent transmission to third-party models or helper processes is a meaningful privacy and compliance risk.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The helper exposes broad arbitrary process-execution capability that is not constrained to PDF operations, despite the skill being described as PDF-focused. If higher-level code passes user-influenced commands or environment variables into this wrapper, the skill could be used to execute unintended programs, expand attack surface, and bypass the intended functional scope.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · pdfagent/core/exec.py (reported line 21)May include surrounding context.

python
env: dict[str, str] | None = None,
) -> str:
    try:
        result = subprocess.run(
            list(cmd),
            cwd=str(cwd) if cwd else None,
            input=input_text,

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This file adds generic LLM orchestration and external command execution capability to a skill described as PDF-only operations, materially expanding its attack surface beyond the stated scope. The code allows operator-supplied commands and arbitrary prompt forwarding to subprocesses, which can enable unintended data exfiltration or execution of non-PDF-related tooling if this feature is exposed or misconfigured.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Prompts are sent directly to external commands or to ollama subprocesses without any disclosure, consent, or visible policy enforcement in this file. If prompts contain PDF-derived sensitive content, this behavior can leak confidential data to external local services, wrappers, logs, or downstream model infrastructure, especially when the command provider is configurable.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The function invokes Ghostscript via a subprocess at L42 and writes output files at L29 and L54, including overwriting the PDF when metadata stripping is enabled. The code contains no confirmation prompt, print/log message, or explanatory docstring/comment warning users about these filesystem and command-execution side effects.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
78% confidence
Finding

The code locates and invokes wkhtmltopdf, Chrome, or Chromium via run_cmd for conversion. Spawning external executables is a powerful host capability and is not clearly disclosed by the manifest's brief description of PDF operations, especially for a self-hosted skill where capability boundaries matter.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The function invokes external binaries (wkhtmltopdf and later Chrome/Chromium) via run_cmd to process input into a PDF. Although this may be part of the implementation, the file contains no print/log statement, prompt, or explanatory comment/docstring disclosing that external commands will be executed.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The fallback loader accepts arbitrary http/https URLs and fetches them server-side with urlopen, which gives this skill outbound network access beyond purely local PDF conversion. In a self-hosted conversion tool, this can be abused for SSRF-style access to internal services or unintended data egress, especially because no allowlist, timeout, or disclosure is present.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

Remote HTML retrieval occurs silently, so users may believe they are performing only local/self-hosted conversion while the tool makes outbound requests. That lack of transparency increases the risk of accidental data disclosure, policy violations, and misuse of the host as a network pivot.

Content

No source excerpt is available for this finding.

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill allows unbounded resource consumption (API calls, storage, compute). Without rate limits or quotas, a compromised or misbehaving agent can cause denial-of-service or cost overruns.

Content

Scanner excerpt · pdfagent/tools/office.py (reported line 11)May include surrounding context.

python
def office_to_pdf(input_path: Path, output_dir: Path, timeout_sec: int = 60) -> Path:
    env_timeout = None
    try:
        env_timeout = int(os.environ.get("PDFAGENT_SOFFICE_TIMEOUT", "0"))
    except ValueError:

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill allows unbounded resource consumption (API calls, storage, compute). Without rate limits or quotas, a compromised or misbehaving agent can cause denial-of-service or cost overruns.

Content

Scanner excerpt · pdfagent/tools/office.py (reported line 15)May include surrounding context.

python
def office_to_pdf(input_path: Path, output_dir: Path, timeout_sec: int = 60) -> Path:
    env_timeout = None
    try:
        env_timeout = int(os.environ.get("PDFAGENT_SOFFICE_TIMEOUT", "0"))
    except ValueError:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

In rasterize mode, the code renders the source PDF into PNG files inside a predictable temporary directory under the output path and leaves those intermediate artifacts on disk. Because redaction is commonly used on sensitive documents, these unredacted page images can expose the very content the user intended to remove, especially on shared hosts, multi-user systems, or when working directories are backed up, synced, or later inspected.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.