Back to skill

Security audit

pdf-processor

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent local PDF toolkit, but it needs review because it exposes PDF passwords and can generate unsafe spreadsheet content from untrusted PDFs.

Review this before installing if you will handle sensitive PDFs. Do not pass real passwords on the command line or rely on encrypt_pdf.py without removing the password printout. Treat generated XLSX files from untrusted PDFs as potentially active content, use separate output directories and backups for batch operations, and install dependencies in an isolated environment with pinned versions.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/pdf_to_excel.py:34
Finding

Spreadsheet Formula Injection in Generated Excel Workbooks

Content
View full analysis
1 else [] df = pd.DataFrame(data, columns=headers) df.to_excel(writer, sheet_name=sheet_name[:31], index=False) ``` From `scripts/extract_tables.py`: ```python with pd.ExcelWriter(output_path, engine='openpyxl') as writer: for idx, table_info in enumerate(tables_data): sheet_name = f"Page{table_info['page']}_Table{table_info['table_index']}" df = pd.DataFrame(table_info['data'][1:], columns=table_info['data'][0]) df.to_excel(writer, sheet_name=sheet_name[:31], index=False) ``` ### Technical Analysis The scripts treat PDF text, table headers, and table cells as trusted data and write them directly to XLSX workbooks through Pandas and OpenPyXL. An attacker can construct a PDF whose extracted cell content begins with a formula marker, particularly `=`. Such values may be stored as spreadsheet formulas rather than inert text. The vulnerability affects both table content and, depending on serialization behavior, column headers derived from the PDF. No validation or neutralization occurs before workbook generation. Formula behavior depends on the spreadsheet application and its security configuration. Potential payloads include external workbook references, network-triggering formulas, deceptive hyperlinks, and application-specific formula mechanisms capable of exposing data. ### Attack Path 1. An attacker creates a PDF containing a table cell such as an external-reference or hyperlink formula beginning wit ...[truncated 1105 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/encrypt_pdf.py:30
Finding

Encryption Password Exposed Through Command-Line Arguments and Console Output

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/pdf_to_word.py:51
Finding

Predictable Temporary Image Files Permit File Overwrite and Unsafe Deletion

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
requirements.txt:5
Finding

Non-Reproducible Dependency Installation Uses Unbounded Versions Without Integrity Hashes

Content
View full analysis
=1.23.0 pdfplumber>=0.10.0 # Word/Excel 转换 python-docx>=1.1.0 openpyxl>=3.1.0 # 图片处理 Pillow>=10.0.0 # 可选: OCR 支持 # pytesseract>=0.3.10 # tessdata>=4.1.0 # 测试 pytest>=7.4.0 pytest-cov>=4.1.0 ``` From `SKILL.md`: ```bash pip install -r requirements.txt ``` ```bash pip install pymupdf pdfplumber python-docx openpyxl pillow ``` ### Technical Analysis All active dependencies use minimum-version constraints rather than reviewed exact versions. No hashes or lock file are provided. Consequently, two installations performed at different times can retrieve and execute different package releases. Python package installation may execute package build or installation logic. If an upstream package, account, or distribution channel is compromised, an unreviewed release satisfying the broad constraint can enter the environment without any project change. The code also imports Pandas in `pdf_to_excel.py`, while `pandas` is absent from `requirements.txt`. This mismatch encourages ad hoc installation and further reduces reproducibility. No evidence of typosquatted package names, custom package indexes, or an actually malicious dependency was found. This finding concerns dependency integrity and version-control weaknesses rather than a confirmed compromised package. ### Attack Path 1. A user follows the documentation and runs `pip install -r requirements.txt`. 2. Pip resolves the newest available packages satisfying the `>=` constraints. 3. A newly released, compromised, or otherwise unsafe package version satisfies those constraints. 4. Because no expected hashes or exact reviewed versions are specified, installation proceeds. 5. Package installation or later import executes t ...[truncated 751 chars]
Remediation
View remediation
--hash=sha256: pdfplumber== --hash=sha256: pandas== --hash=sha256: python-docx== --hash=sha256: openpyxl== --hash=sha256: Pillow== --hash=sha256: ``` ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (40)

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

描述宣称的是功能非常全面的 PDF 处理套件,而提供的代码片段实际是一个批量处理入口脚本,只支持 extract_text、extract_images、add_watermark 和 compress 四种操作。虽然这些功能属于声明范围的一部分,且“批量处理”这一点与描述一致,但大量核心声明能力在代码中没有体现,包括 Word/Excel 转换、合并拆分、OCR、表格提取、加密解密等。因此描述对该代码片段的能力范围存在明显夸大,属于描述与实际行为不一致。代码未显示任何额外的、未声明的敏感能力;问题主要在于过度声明功能。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The description presents a multi-function PDF toolkit, while the actual code shown implements only one specific feature: decrypting an encrypted PDF given the correct password. Although decryption is mentioned in the declared description ('加密解密'), the overall declared purpose materially overstates what this supplied code chunk actually does. There are no undeclared risky behaviors or inconsistent resource accesses; the mismatch is that the code's real scope is much narrower than the declared one.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

该代码块的主要功能非常单一:使用 PyMuPDF 遍历 PDF 页面,提取其中图片并保存到本地目录,同时输出索引文件。虽然“图片提取”属于声明能力中的一个子项,但声明将该技能描述为覆盖大量 PDF 处理场景的一站式工具,而当前提供的代码并未体现这些核心能力中的绝大多数。因此这是明显的描述与实际行为不一致,属于显著夸大声明范围,而不是仅仅实现细节未展示。代码本身也没有出现额外未声明的敏感能力;问题在于声明远超实际实现。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

该代码块的行为很明确:使用 pdfplumber 遍历 PDF 页面并提取表格,随后可借助 pandas/openpyxl 将表格分别写入 Excel 工作表。它没有实现声明中列出的多数功能,也没有体现“一站式 PDF 处理技能”的综合能力。虽然“表格提取”属于声明范围的一部分,但声明将技能描述为多功能 PDF 工具,而当前代码只是其中一个非常具体的子功能,因此描述不能准确代表该代码块的实际行为,构成描述与行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的是一个功能全面的 PDF 处理技能,但当前代码块的实际行为非常有限,仅执行 PDF 文本提取,并可选提取元数据与保存输出。其主要用途只覆盖声明中的“从 PDF 提取文本内容”这一小部分,无法支持描述中列出的大多数核心功能。因此描述与该代码块实际能力存在明显不匹配。虽然元数据提取属于较小的附加能力,不算严重风险,但整体上该代码块远不足以支撑“一站式 PDF 处理”这一声明。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

声明描述的是一个覆盖面很广的综合 PDF 处理技能,但提供的代码块功能非常单一,只做 OCR 识别。它打开输入 PDF,将每页渲染为图像,用 pytesseract 提取文字,然后生成新的 PDF,把原页面图像插入并在页面下方附加 OCR 文本。代码没有看到任何与 Word/Excel 转换、合并拆分、批量处理、水印、加密、压缩、表格提取等相关的实现。因此描述显著夸大了能力范围,属于描述与实际行为不符。虽然 OCR 属于声明的一部分,但不足以支撑“一站式 PDF 处理技能”的整体表述。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

该代码块的实际功能是“PDF 转 Excel(附带文本与表格提取)”,属于声明能力集合中的一个子集。虽然代码行为本身没有出现危险的未声明额外能力,也没有访问与声明不一致的资源,但声明将该技能描述为覆盖广泛 PDF 工具链的一站式方案,而当前提供的代码仅支持单文件 PDF 到 Excel 的转换及文本/表格提取,主用途明显比声明窄得多。就‘描述是否准确代表该代码实际行为’而言,存在显著不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

该描述将技能定义为“ 一站式 PDF 处理 ”,覆盖提取、转换、合并拆分、OCR、批处理、水印、加密、压缩等广泛能力;但提供的代码块仅是一个 PDF 拆分脚本,使用 PyMuPDF 打开单个 PDF,并按页或按给定页码范围输出拆分文件。它没有实现文本提取、格式转换、合并、OCR、批量处理、水印、加密/解密或压缩等功能。虽然“拆分 PDF”属于声明中的一部分,但整体描述显著夸大了代码实际能力范围,因此描述与行为不准确,构成不匹配。

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The skill documents decryption, encryption, batch processing, splitting, merging, compression, and output-generating commands without warning about irreversible overwrites, bulk modification, or safe output handling. In a document-processing skill, that omission increases the chance that users or agents will modify sensitive files in place, decrypt protected content, or mass-process the wrong directory without adequate confirmation.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SKILL.md (reported line 137)May include surrounding context.

安装 Tesseract OCR

Windows: https://github.com/UB-Mannheim/tesseract/wiki

macOS: brew install tesseract

Linux: sudo apt install tesseract-ocr

pip install pytesseract

text

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code file contains natural-language instructions, usage text, and argument help entirely in Chinese, which imposes a specific language on users. The file does not offer a language choice or explain that the skill is intentionally region-specific, matching the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The script accepts the PDF password as a positional command-line argument, which can expose the secret through process listings, shell history, audit logs, and job-control tooling on multi-user or monitored systems. Because this skill explicitly handles PDF decryption, the secret is highly sensitive and the context makes the leakage more relevant rather than less.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The script prints the PDF encryption password directly to standard output after completing encryption. This can leak the secret into terminal history, logs, CI job output, or agent transcripts, undermining the protection provided by the PDF encryption feature. In the context of an automation skill, this is more dangerous because outputs are often captured and persisted by surrounding systems.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file header, usage text, and subsequent user-facing strings indicate the skill is designed to communicate only in Chinese. Under the policy, forcing a specific language without opt-in or documented justification is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This code file includes its title, usage/help text, docstrings, and status messages in Chinese, which effectively forces a specific language for interaction. The policy allows locale constraints only when users are given a choice or when the restriction is clearly justified as region-specific, neither of which is present here.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code file contains natural-language documentation and console output entirely in Chinese, including the top-level usage instructions. Under the policy, forcing a specific language without user opt-in is a locale policy violation unless the constraint is clearly documented and justified, which is not present here.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest description and the rest of the skill documentation are entirely in Chinese, and the OCR section specifically emphasizes Chinese recognition. There is no indication that the skill supports user-selected language/locale preferences or that it is intentionally limited to a Chinese-only regional context.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The natural-language comments in this requirements file are entirely in Chinese, which imposes a specific language on users without offering any language choice or explaining a region-specific need. Under the policy for natural-language violations, a forced locale or language without opt-in can be considered a policy issue.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
93% confidence
Finding

Using a lower-bounded but unpinned dependency for PyMuPDF allows future installs to resolve to different versions over time, including versions with newly introduced vulnerabilities or breaking security changes. In a PDF-processing skill that handles untrusted documents, dependency drift materially increases supply-chain risk and makes vulnerability exposure hard to assess or reproduce.

Content

Scanner excerpt · requirements.txt (reported line 5)May include surrounding context.

text
# 安装: pip install -r requirements.txt

# 核心 PDF 处理
pymupdf>=1.23.0
pdfplumber>=0.10.0

# Word/Excel 转换

Unverifiable Dependency: pymupdf has 2 known advisory(ies) (CVE-2026-3029 (PyMuPDF has a path traversal in _main_.py); CVE-2026-3029 (PyMuPDF has a path traversal in _main_.py)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
90% confidence
Finding

PyMuPDF has known advisories, and because the manifest does not pin an exact version, there is no way to verify whether installations will receive a fixed or vulnerable release. This uncertainty is more serious in a PDF-processing skill because the library is directly exposed to untrusted file content and may include auxiliary command-line paths with file handling risks.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
88% confidence
Finding

An unpinned pdfplumber dependency means builds are not reproducible and may silently pull in newer releases with security defects or incompatible parser behavior. Because this skill parses attacker-controlled PDF content, unexpected changes in the parsing stack can increase exposure to denial-of-service or parser-related vulnerabilities.

Content

Scanner excerpt · requirements.txt (reported line 6)May include surrounding context.

text
# 核心 PDF 处理
pymupdf>=1.23.0
pdfplumber>=0.10.0

# Word/Excel 转换
python-docx>=1.1.0

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
91% confidence
Finding

Unpinned python-docx introduces supply-chain uncertainty and makes it impossible to know which code will be installed in different environments. Since the skill performs format conversion and may ingest complex document structures, dependency drift can reintroduce historical XML parsing flaws or other unsafe behaviors.

Content

Scanner excerpt · requirements.txt (reported line 9)May include surrounding context.

text
pdfplumber>=0.10.0

# Word/Excel 转换
python-docx>=1.1.0
openpyxl>=3.1.0

# 图片处理

Unverifiable Dependency: python-docx has 2 known advisory(ies) (CVE-2016-5851 (Improper Restriction of XML External Entity Reference in python-docx); CVE-2016-5851 (python-docx before 0.8.6 allows context-dependent attackers to conduct XML Exter)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
89% confidence
Finding

python-docx has historical XXE-related advisories, and the unpinned requirement prevents confirming whether deployed environments use a patched release. In a document conversion workflow, unsafe XML handling can expose local files or internal network resources if malicious documents are processed by a vulnerable version.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
91% confidence
Finding

Leaving openpyxl unpinned permits uncontrolled upgrades and hampers reproducibility, incident response, and CVE triage. For a skill converting data to or from spreadsheet formats, this is more concerning because document parsers often process attacker-controlled structured content and have a history of XML-related issues.

Content

Scanner excerpt · requirements.txt (reported line 10)May include surrounding context.

text
# Word/Excel 转换
python-docx>=1.1.0
openpyxl>=3.1.0

# 图片处理
Pillow>=10.0.0

Static analysis

No suspicious patterns detected.