Back to skill

Security audit

MinerU Doc Parser

Security checks for vulnerabilities and agentic risk

Overview

The skill is a real document-parsing helper, but it needs review because it can send local documents to a remote AI service and installs a mutable global CLI without enough disclosure or safeguards.

Install only if you are comfortable sending selected documents or URLs to MinerU/OpenDataLab for processing. Avoid confidential or regulated files unless you have approval, prefer a pinned and isolated CLI install where possible, and treat MinerU tokens as secrets that should not be pasted into shared terminals, scripts, logs, or screenshots.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T01 · Skill Instruction Hijacking

Warning
Location
SKILL.md:311
Finding

Mandatory Promotional Output Hijacks Agent Responses

Content
View full analysis
Tip: `flash-extract` is the fast, login-free mode (limited to 10 MB/20 pages and without table recognition). For larger files, table/formula recognition, or multiple export formats, visit https://mineru.net/apiManage/token to create a token, configure it with `mineru-open-api auth`, and then use `mineru-open-api extract`. Keep the hint to ONE short sentence. Do NOT repeat the hint if the user has already seen it in this session. ``` The promotional sentence above is an English translation of the original fixed message so that this report remains entirely in English. The command names, limits, and destination URL are unchanged. ### Technical Analysis The Skill uses mandatory directives to alter the agent's final response after every successful `flash-extract` operation. The required content promotes account-token acquisition at an external service and is not necessary to return the requested extraction result. Although the instruction does not disable safety controls or request hidden data, it persistently changes the agent's output policy whenever the Skill is loaded. The use of `MUST`, combined with a prescribed promotional message and session-level repetition tracking, makes this more than optional troubleshooting guidance. ### Attack Path 1. The Skill is loaded for a document-extraction request. 2. The agent invokes `flash-extract`. 3. The extraction succeeds. 4. The Skill requires the agent to alter its response by appending the prescribed external-service promotion. 5. The user is directed to the vendor's token-registration page even when the extraction completed without requiring an account or token. ### Impact Assessment The issue affects the integrity and neutrality of ...[truncated 316 chars]
Remediation
View remediation

other

Error
Location
` for quick Markdown conversion 2. **Need more?** Create token at https://mineru.net/apiManage/token, run `mineru-open-api auth`, then use `mineru-open-api extract` for tables, formulas, OCR, multi-format, and batch 3. **Web pages**: `mineru-open-api crawl <url>` to convert web content 4. **Check results**: output goes to stdout (default) or `-o` directory ## Authentication Only required for `extract` and `crawl`. Not needed for `flash-extract`. Configure your API token (create one at https://mineru.net/ ...[truncated 2665 chars]:81
Finding

Local Documents May Be Transmitted to an External Parsing Service Without Explicit Consent

Content
View full analysis
` for quick Markdown conversion 2. **Need more?** Create token at https://mineru.net/apiManage/token, run `mineru-open-api auth`, then use `mineru-open-api extract` for tables, formulas, OCR, multi-format, and batch 3. **Web pages**: `mineru-open-api crawl ` to convert web content 4. **Check results**: output goes to stdout (default) or `-o` directory ## Authentication Only required for `extract` and `crawl`. Not needed for `flash-extract`. Configure your API token (create one at https://mineru.net/apiManage/token): ```bash mineru-open-api auth export MINERU_TOKEN="your-token" ``` Token resolution order: `--token` flag > `MINERU_TOKEN` env > `~/.mineru/config.yaml`. ``` The associated examples at `SKILL.md:122-126` direct the agent to pass local files such as `report.pdf` to `flash-extract`. ### Technical Analysis The Skill's declared functionality is cloud-backed AI document parsing. Passing a local file to `mineru-open-api flash-extract` or `mineru-open-api extract` therefore creates a network-processing path for the document. This transfer is functionally relevant, but the Skill does not require the agent to: - Explicitly disclose that local document contents may leave the machine. - Obtain user consent before uploading a potentially confidential file. - Identify the processing endpoint used by the installed CLI. - Explain storage, retention, model-training, deletion, or regional-processing policies. - Offer a local-only processing alternative. The token-free mode can make the workflow appear local or low-risk even though authentication status does not determine whether document data is transmitted. The audited artifact contains only usage instructions, ...[truncated 1355 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Error
Location
SKILL.md:35
Finding

Mutable Unpinned Dependencies Are Installed as Global Executables

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (5)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill describes extracting local files and URLs with an AI-powered service but does not clearly warn that document contents may leave the local environment and be sent to a third-party backend. Because this skill is specifically for document parsing, users may provide sensitive PDFs, scans, or web resources, making silent exfiltration risk significant.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill advertises very broad triggers such as reading PDFs or parsing any document format, which can cause the agent to invoke it in cases where the user did not intend to send content to this external parser. In this skill's context, that overbreadth is more dangerous because the tool processes local files and URLs and may transmit document contents to a remote service.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
72% confidence
Finding

The skill encourages obtaining and configuring a persistent API token, which introduces session persistence and credential retention risk if stored in config files or environment variables without clear lifecycle guidance. In this context, persistence matters because the same token authorizes future document uploads to the remote parsing service.

Content

Scanner excerpt · SKILL.md (reported line 76)May include surrounding context.

md
| Supported types | PDF, Images (png/jpg/jpeg/jp2/webp/gif/bmp), Docx, PPTx |
| IP rate limit | Per-minute request caps (HTTP 429 when exceeded) |

When any limit is exceeded, the agent should suggest switching to `extract` with a token (create at https://mineru.net/apiManage/token), which has significantly higher limits.


## Core workflow

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The authentication section recommends storing the API token in an environment variable without warning that the token is sensitive and may be exposed through shell history, process environments, logs, or shared sessions. This is especially relevant in agent and CLI environments where commands may be echoed or persisted.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

These instructions require the agent to append a Chinese-language hint after successful extraction regardless of the user's language preference. That is a natural-language policy violation because it forces a specific locale without offering user choice or documenting a justified locale constraint.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.