Back to skill

Security audit

PDF Catalog

Security checks for vulnerabilities and agentic risk

Overview

This is a local PDF catalog extraction skill with disclosed file generation and optional Excel filling, but users should treat dependency installation, workbook edits, and generated Markdown as areas needing care.

Install in an isolated virtual environment, pin dependency versions where possible, run the script only on trusted PDF folders, and pass --excel-path only for a workbook you intend to modify, preferably a copy. Treat generated Markdown as extracted data, not trusted agent instructions, especially when PDF filenames or source files are supplied by others.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:295
Finding

Unpinned Third-Party Dependencies Create a Supply-Chain Risk

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:295 and README.md:9
Vulnerability Type: Unpinned dependency installation
Risk Level: Medium

Vulnerable Code

SKILL.md:295:

bash
pip3 install docling openpyxl

README.md:9:

bash
pip3 install docling openpyxl

Technical Analysis

The installation instructions retrieve the latest available versions of docling and openpyxl without version constraints, package hashes, a lock file, or an explicitly trusted package index. Consequently, the code installed by users can change after this Skill has been reviewed.

Python packages can execute code during installation and whenever imported. The script imports both dependencies near startup, meaning a compromised or unexpectedly replaced package version could execute with the same privileges as the user running the Skill.

The repository does not itself contain a malicious dependency, and no direct dependency-confusion package name was identified. The vulnerability is the absence of controls that guarantee installation of the reviewed dependency artifacts.

Attack Path

  1. An upstream package release or one of its transitive dependencies is compromised.
  2. A user follows the documented command and installs the latest packages from the configured Python package index.
  3. The package manager downloads the compromised release because no version or hash constraint prevents it.
  4. Malicious code executes during package installation or when extract.py imports the dependency.
  5. The malicious code operates with the privileges and data access of the installing or executing user.

Impact Assessment

Successful exploitation could provide arbitrary code execution under the current user's account. Depending on that user's privileges, the attacker could access local files, PDF source material, output data, Excel workbooks, environment variables, and credentials available to the process ...[truncated 126 chars]

Remediation
View remediation

Remediation Suggestions

  • Pin every direct dependency to a reviewed version, for example through a version-controlled requirements or lock file.
  • Pin transitive dependencies as well, using a tool such as pip-tools, Poetry, or an equivalent reproducible dependency manager.
  • Record and verify package hashes, then install with pip install --require-hashes.
  • Configure an explicitly trusted package index rather than relying on an unspecified environment configuration.
  • Run dependency vulnerability and provenance checks in CI.
  • Periodically update dependencies through reviewed changes rather than automatically consuming the newest release.
  • Install and execute the Skill in an isolated virtual environment or container with access limited to required input and output paths.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/extract.py:237
Finding

Unsanitized PDF Filenames Are Embedded in Generated Markdown

Content
View full analysis

Vulnerability Details

File Location: scripts/extract.py:237-260
Vulnerability Type: Persistent Markdown content injection
Risk Level: Medium

Vulnerable Code

python
for p in products:
    model = p['model_no'] or '❌ 未找到'
    pkg = ', '.join(p['package_specs']) if p['package_specs'] else 'N/A'
    items = ', '.join(p['customer_items'][:3]) if p['customer_items'] else '❌ 未找到'
    lengths = ', '.join(p['lengths'][:3]) if p['lengths'] else 'N/A'

    f.write(f"| {p['pdf_number']} | {p['pdf_file']} | {model} | {pkg} | {items} | {lengths} |\n")

f.write("\n---\n\n## 详细数据\n\n")
f.write("完整的产品数据已保存在:`产品详细数据.json`\n")


def generate_product_cards(products, output_dir):
    """生成单个产品词条"""
    cards_dir = os.path.join(output_dir, '类目词条')
    os.makedirs(cards_dir, exist_ok=True)

    for p in products:
        pdf_base = p['pdf_file'].replace('.pdf', '')
        card_path = os.path.join(cards_dir, f"{pdf_base}-类目词条.md")

        with open(card_path, 'w', encoding='utf-8') as f:
            f.write(f"# {p['pdf_file']} 产品类目词条\n\n")

Technical Analysis

p['pdf_file'] originates from names returned by os.listdir() and is used without validation or Markdown escaping. On filesystems that permit control characters in filenames, a crafted PDF name can contain newlines, table delimiters, Markdown links, headings, or instruction-like text.

The value is inserted directly into the generated index and product-card Markdown. As a result, an attacker-controlled filename can escape the intended table cell or heading context and introduce persistent content into the generated knowledge base.

The filename is also used to construct a product-card filename. Because os.listdir() returns a directory entry name rather than a path containing separators, this code does not by itself establish a conventional slash-based directory traversal path. The demonstrated issue is Markdown and knowledge-base conten ...[truncated 1218 chars]

Remediation
View remediation

Remediation Suggestions

  • Validate input filenames against a strict allowlist such as letters, digits, periods, underscores, and hyphens.
  • Reject filenames containing control characters, carriage returns, newlines, null bytes, or Unicode line separators.
  • Escape Markdown-special characters according to the output context. Table cells should escape pipes and normalize line breaks; headings should remove or escape structural syntax.
  • Generate output filenames from a sanitized identifier rather than directly from the source filename.
  • Verify generated paths with os.path.realpath() and ensure they remain beneath the intended output directory.
  • Store original filenames as JSON data when exact preservation is required, while using an encoded or sanitized display value in Markdown.
  • Treat generated knowledge-base files as untrusted data when they are supplied to an AI Agent, and separate extracted content from Agent instructions.
  • Add tests using filenames containing newlines, pipes, brackets, Markdown links, headings, and other control or formatting characters.
Vulnerability Patterns
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (7)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The entire skill documentation is written in Chinese, including headings, instructions, and output descriptions, with no indication that another language is supported or that Chinese is a required locale for a region-specific use case. This can violate language/locale policy when a skill implicitly requires a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill documents shell execution and file read/write behavior via pdftotext, pdftoppm, OCR processing, JSON/Markdown generation, and Excel modification, but it does not declare any tool scope such as allowed-tools or permissions. That omission weakens least-privilege controls and can let an agent invoke broader filesystem or shell capabilities than reviewers and operators expect.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This code file contains natural-language strings such as the module docstring, warnings, and CLI help text that assume Chinese as the required language. Under the stated policy, forcing a specific language without user opt-in is a natural-language policy violation unless the locale constraint is explicitly justified.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/extract.py (reported line 68)May include surrounding context.

python
"""
    # Method 1: pdftotext (矢量图 PDF)
    try:
        result = subprocess.run(['pdftotext', pdf_path, '-'], capture_output=True, text=True, timeout=30)
        text = result.stdout
        
        if len(text) >= ocr_threshold and not use_ocr:

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The manifest describes extraction from PDF catalogs and generation of structured outputs plus Excel fill data, which implies producing data for Excel use. This function goes further by opening a user-supplied workbook and saving changes back to the same file, which is a state-changing operation not clearly conveyed by the skill description.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The Excel fill step writes model numbers into workbook cells, but the skill description does not clearly warn users that the specified Excel file will be modified. This can cause unintended data changes, especially when operators expect analysis-only behavior or run the skill against important production spreadsheets.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

This file presents all user-facing content in a single language, which may violate a language/locale policy when no user opt-in or justification is provided. The content does not indicate that the skill is intended only for Chinese-speaking users or a region-specific workflow.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.