Back to skill

Security audit

Pdf Extractor Skill

Security checks for vulnerabilities and agentic risk

Overview

This PDF extraction skill is mostly purpose-aligned, but its optional LLM mode can send PDF content to remote services and includes a hardcoded fallback API key.

Review this skill before installing. Use local-only conversion for sensitive PDFs unless you intentionally want LLM enhancement, understand which remote endpoint will receive document-derived content, and provide your own scoped API key. The embedded Volcengine key should be removed or rotated by the publisher, and dependency setup should ideally be pinned or run in an isolated environment.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/pdf2md_marker.py:29
Finding

Hardcoded Volcengine API Credential

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
README.md:15
Finding

Unpinned Third-Party Dependency Installation

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (10)

Ssd 3

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The embedded fallback API key creates a path where PDF contents can be sent to an external LLM service using credentials bundled in the script, rather than credentials deliberately provided by the user. Because this tool handles document extraction, the context makes accidental disclosure of confidential PDFs especially dangerous.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

A hardcoded fallback API key allows the script to authenticate to an external LLM service even when the operator did not intentionally configure credentials. If LLM mode is enabled, sensitive PDF content may be sent to a remote endpoint under an embedded key with inadequate notice, creating both data leakage and credential misuse risk.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 7)May include surrounding context.

Installation

1. Create Conda Environment

bash
conda create -n pdf-extractor python=3.10 -y

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The README advertises an optional LLM enhancement flag but does not disclose that document content may be transmitted to an external service for processing. Because academic PDFs can contain unpublished research, personal data, or confidential material, this omission can lead users to exfiltrate sensitive content without informed consent.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill says it is self-contained and should not perform extra installs or downloads, yet later states that Marker models are downloaded automatically. This creates a misleading trust boundary: using the skill may trigger unexpected network access and retrieval of remote model artifacts, which can violate offline-use assumptions, enterprise policy, or reproducibility expectations.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documented --ark-code-latest and custom OpenAI-compatible parameters imply that PDF content may be transmitted to an external service, but the skill does not warn users about this data flow. Because PDFs may contain sensitive unpublished research, proprietary data, or personal information, omission of this disclosure can lead to unintended exfiltration and policy violations.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This code sends the user prompt and any attached image data to an external OpenAI-compatible HTTP service at a configured base URL, which constitutes data exfiltration to another service boundary. Even though the default endpoint is localhost, the code provides no consent, classification, or safeguards around potentially sensitive prompt/image contents, so deployments that point this URL to a remote proxy could unintentionally disclose confidential data.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/marker_openai_compat_service.py (reported line 38)May include surrounding context.

python
openai_base_url: Annotated[
        str, "The base url to use for OpenAI-like models. No trailing slash."
    ] = "https://api.openai.com/v1"
    openai_model: Annotated[str, "The model name to use for OpenAI-like model."] = (
        "gpt-4o-mini"
    )

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The script adds optional outbound LLM/network capability to a PDF-to-Markdown converter and pairs it with a built-in credential fallback, which can cause document contents to be transmitted to a third-party service during conversion. In this context the skill processes potentially sensitive PDFs, so enabling remote processing without strong gating, explicit consent, or secret management materially increases data-exfiltration risk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The function assigns a default language list of ['Chinese', 'English'] when the caller does not specify languages. This imposes a locale/language assumption by default rather than offering a neutral default or explicit opt-in, which matches the policy concern about forcing a specific language setting.

Content

No source excerpt is available for this finding.