Back to skill

Security audit

Pdf Field Extractor

Security checks for vulnerabilities and agentic risk

Overview

The skill matches its PDF extraction purpose, but it can send sensitive document text and an API key to a configurable outside AI endpoint without strong safeguards.

Install only if you are comfortable processing sensitive PDFs through the model endpoint you configure. Use a trusted HTTPS endpoint, avoid custom or untrusted api_base values, do not reuse an API key with unrelated providers, review generated spreadsheets before sharing, and avoid identity, banking, or confidential business documents unless your data-handling requirements allow remote AI processing.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/field_extractor.py:253
Finding

Unrestricted API endpoint can receive sensitive document content and bearer credentials

Content
View full analysis

Vulnerability Details

File Location: scripts/field_extractor.py:253-279
Vulnerability Type: Unrestricted sensitive-data transmission and server-side request forgery exposure
Risk Level: High

Vulnerable Code

python
api_base = api_base or DEFAULT_API_BASE
model = model or DEFAULT_MODEL

# Build messages
system_prompt = SYSTEM_PROMPTS.get(doc_type, SYSTEM_PROMPTS["generic"])
user_prompt = build_user_prompt(doc_type, custom_fields).format(text=text[:8000])  # Truncate to 8k chars

messages = [
    {"role": "system", "content": system_prompt},
    {"role": "user", "content": user_prompt},
]

# Call API
endpoint = f"{api_base.rstrip('/')}/chat/completions"
headers = {
    "Content-Type": "application/json",
    "Authorization": f"Bearer {api_key}",
}
payload = {
    "model": model,
    "messages": messages,
    "temperature": temperature,
    "max_tokens": 2048,
}

try:
    response = requests.post(endpoint, headers=headers, json=payload, timeout=timeout)

Technical Analysis

The AI-assisted extraction feature legitimately requires sending document text to a model provider. However, the caller-controlled api_base is used without validating its URL scheme, hostname, resolved IP address, or trust relationship with the supplied credential.

The resulting request transmits both:

  • Up to 8,000 characters of PDF-derived content, potentially including identity numbers, bank account details, addresses, contractual terms, invoices, and other confidential information.
  • The API key in an Authorization: Bearer header.

An endpoint using plaintext HTTP can expose both values to network interception. A malicious endpoint can directly collect them. Depending on network accessibility, loopback, private, link-local, and cloud metadata destinations may also be reachable, creating server-side request forgery exposure.

A timeout limits request duration but does not restrict de ...[truncated 1572 chars]

Remediation
View remediation

Remediation Suggestions

  1. Require HTTPS and reject plaintext HTTP endpoints.
  2. Use a provider-specific allowlist for model API hostnames by default.
  3. Make custom endpoints an explicit opt-in accompanied by a warning that document content and credentials will be transmitted.
  4. Resolve the destination hostname and reject loopback, private, link-local, multicast, and cloud metadata address ranges unless an administrator explicitly authorizes them.
  5. Revalidate every redirect destination or disable redirects for API requests.
  6. Bind each API credential to its expected provider hostname and refuse to forward it to a different host.
  7. Display or log the destination hostname before transmission without logging the API key or document content.
  8. Require explicit consent before transmitting sensitive document classes such as identity documents and bank statements.
  9. Consider optional local extraction or redaction of sensitive fields before remote processing.
  10. Add tests covering HTTP rejection, malicious redirects, DNS resolution to private addresses, metadata endpoints, and provider/credential mismatches.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/output_generator.py:68
Finding

Untrusted PDF and model output can produce executable spreadsheet formulas

Content
View full analysis

Vulnerability Details

File Location: scripts/output_generator.py:68-121
Vulnerability Type: Spreadsheet formula injection
Risk Level: Medium

Vulnerable Code

python
# Collect all unique field keys across all results
all_keys = set()
for result in results:
    all_keys.update(result.keys())

# Filter out internal fields
internal_fields = {"_filename", "_timestamp", "_doc_type", "_page_count", "_is_scanned"}
display_keys = sorted([k for k in all_keys if k not in internal_fields])

# Build column headers
if include_metadata:
    headers = ["文件名", "提取时间", "文档类型"] + display_keys
else:
    headers = display_keys

# Create workbook
wb = openpyxl.Workbook()
ws = wb.active
ws.title = sheet_name

# Write headers
for col_idx, header in enumerate(headers, start=1):
    cell = ws.cell(row=1, column=col_idx, value=header)
    cell.font = HEADER_FONT
    cell.fill = HEADER_FILL
    cell.alignment = HEADER_ALIGNMENT
    cell.border = CELL_BORDER

# Write data rows
for row_idx, result in enumerate(results, start=2):
    if include_metadata:
        ws.cell(row=row_idx, column=1, value=result.get("_filename", "")).font = CELL_FONT
        ws.cell(row=row_idx, column=1).border = CELL_BORDER
        ws.cell(row=row_idx, column=1).alignment = CELL_ALIGNMENT

        ws.cell(row=row_idx, column=2, value=result.get("_timestamp", "")).font = CELL_FONT
        ws.cell(row=row_idx, column=2).border = CELL_BORDER
        ws.cell(row=row_idx, column=2).alignment = CELL_ALIGNMENT

        doc_type_display = {
            "invoice": "发票",
            "contract": "合同",
            "receipt": "收据",
            "bank_statement": "银行对账单",
            "license": "营业执照",
            "id_card": "身份证/护照",
            "express": "快递单",
            "generic": "通用文档",
        }.get(result.get("_doc_type", ""), result.get("_doc_type", ""))
        ws.cell(row=row_idx, column=3, value=
...[truncated 2797 chars]
Remediation
View remediation

Remediation Suggestions

  1. Create a single spreadsheet-sanitization function and apply it to every dynamic header, filename, metadata value, and extracted field.
  2. Treat strings beginning with =, +, -, @, tab, carriage return, or newline as unsafe.
  3. Prefix unsafe strings with an apostrophe or otherwise force the cell to contain literal text rather than a formula.
  4. Perform sanitization after converting dictionaries and lists to JSON strings.
  5. Do not rely only on model prompts to prevent formula content; model output and source documents must remain untrusted.
  6. Consider offering a strict export mode that stores all dynamic values as text.
  7. Add regression tests for every dangerous prefix in field values, field names, filenames, and metadata.
  8. Reopen generated workbooks during testing and verify that malicious inputs are represented as literal strings and not cells with formula data types.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (24)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly sends extracted PDF text to an OpenAI-compatible external API for field extraction, but the user-facing description does not clearly warn that uploaded document contents may leave the local environment. Because the supported documents include invoices, contracts, bank statements, licenses, and ID cards, this can expose highly sensitive personal, financial, or business data without informed user consent.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
91% confidence
Finding

The example shows transmission of extracted document text to an external endpoint (https://api.openai.com/v1) for AI processing. In this skill context, the data can include PII, financial records, contracts, and identity documents, so external transmission materially increases confidentiality and compliance risk if users are unaware or if data handling requirements are strict.

Content

Scanner excerpt · SKILL.md (reported line 28)May include surrounding context.

md
text=text,
    doc_type="invoice",
    api_key="sk-xxx",           # 用户自己的API Key
    api_base="https://api.openai.com/v1",  # 可配置
    model="gpt-4o",
)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The invoice template requires the model to extract and return JSON using Chinese field names such as 发票号 and 买方, and the contract template does the same with Chinese labels. This imposes a specific language/locale in the skill instructions without any user opt-in or documented justification for a region-specific workflow.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The contract prompt instructs the model to return JSON with Chinese key names like 合同号, 签订日期, and 付款条件. Because the file presents these as general prompt templates rather than a clearly justified region-specific tool, this is a natural-language locale policy concern.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

When an API key is provided, the code sends extracted PDF text to extract_fields, which is described as AI extraction and therefore likely involves transmitting document contents off-box. While the docstring mentions an API key, it does not warn that potentially sensitive PDF text may be sent to an external service, and there is no confirmation, logging, or user-facing disclosure here.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The fallback output dictionary uses Chinese keys such as 文档类型, 页数, and 文本预览 unconditionally. This imposes a specific locale on all users without opt-in or documented regional scope, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The returned batch summary counts failures by checking for an '_error' key in each result dictionary, but process_batch never adds that key to output_results for failed ProcessingResult objects. As a result, failed files can be reported as successful, which can mislead operators into trusting incomplete or missing extraction results and suppress incident visibility.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · references/prompts.md (reported line 98)May include surrounding context.

md
# ─── Default API Configuration ───────────────────────────────────────────────
DEFAULT_API_BASE = "https://api.openai.com/v1"
DEFAULT_MODEL = "gpt-4o"
DEFAULT_TIMEOUT = 60  # seconds

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/field_extractor.py (reported line 17)May include surrounding context.

python
# ─── Default API Configuration ───────────────────────────────────────────────
DEFAULT_API_BASE = "https://api.openai.com/v1"
DEFAULT_MODEL = "gpt-4o"
DEFAULT_TIMEOUT = 60  # seconds

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill hardcodes Chinese field names across multiple document types and instructs the model to return JSON using those exact keys. This imposes a specific language/locale behavior without offering a user choice or documenting a justified region-specific constraint.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The function always includes raw document text in the prompt payload sent to the external service, with no sanitization, classification, or filtering. Given the document categories, this can transfer highly sensitive PII, financial records, contract terms, and government ID data to a remote provider, creating substantial confidentiality and compliance risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The code sends document text to an external LLM API without any built-in consent, warning, redaction, or policy gate. Because the supported document types include invoices, contracts, bank statements, licenses, and IDs, this can expose sensitive personal, financial, and business data to third-party infrastructure unexpectedly.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

This outbound HTTP request is the concrete mechanism by which sensitive document contents are transmitted off-system. External transmission is not inherently unsafe, but in this context it carries confidentiality risk because the payload contains extracted document text from potentially sensitive records.

Content

Scanner excerpt · scripts/field_extractor.py (reported line 279)May include surrounding context.

python
}

    try:
        response = requests.post(endpoint, headers=headers, json=payload, timeout=timeout)
        response.raise_for_status()
    except requests.exceptions.Timeout:
        raise TimeoutError(f"API request timed out after {timeout}s")

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This code emits user-visible spreadsheet headers in Chinese, and similar hard-coded Chinese text appears throughout Feishu and text message outputs. Because the file provides no opt-in, configuration, or documented region-specific justification, it violates the language/locale policy criteria.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The top-level documentation claims the skill 'handles both text-based PDFs and scanned PDFs', which implies functional support for extracting content from scanned documents. In practice, the code only marks a PDF as scanned when little text is found and can render pages as images for later OCR, but it never performs OCR or extracts text from scanned content.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
81% confidence
Finding

This code defines supported OCR languages as fixed locale lists, starting with only English in the free tier and a limited set in higher tiers. Because the skill behavior is constrained by language support without any user opt-in or documented choice mechanism in this file, it constitutes a natural-language locale policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
77% confidence
Finding

The alias and field mappings embed Chinese-language labels and document terminology as first-class supported inputs, which indicates locale-specific behavior. In the absence of an explicit statement that this is a region-specific skill or that users can choose locale behavior, this may violate language/locale policy expectations.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The supported document types and aliases are predominantly Chinese terms, and the identification logic later prioritizes user hints such as "发票". Because this markdown does not state that the skill is China-specific or otherwise limited to Chinese-language documents, it implies a locale-specific behavior without documented opt-in or justification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The pipeline automatically chooses output paths and writes Excel and JSON files via generate_excel and generate_json. Although saving results is part of the pipeline's purpose, this file does not clearly disclose in a user-facing way that local files will be created, especially when default filenames are used.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The code's user-facing natural language is tied to Chinese examples in the docstring ("发票", "合同") and later returns Chinese display names for document types. Because the skill does not provide any opt-in or alternative locale handling, it reflects a fixed language/locale assumption that may violate language-choice policy.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

get_type_display_name returns Chinese labels such as 发票, 合同, and 通用文档 for all users. This is a natural-language locale choice embedded in code, and the file contains no mechanism for user selection or justification for requiring Chinese output.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

The function automatically reads the OPENAI_API_KEY environment variable to authenticate outbound requests. While this is common practice, this file does not include any explicit user-facing notice that the skill accesses local environment credentials, which is one of the warning-worthy operations for code files.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
98% confidence
Finding

The function docstring says it extracts fields from multiple texts "in parallel" and accepts a max_workers parameter, but the implementation simply iterates over texts in a normal for loop and calls extract_fields one at a time. This is an active contradiction between the documented behavior and the actual code, not just an omitted detail.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The function save_page_as_image performs a file write to output_path, which can overwrite or create files, but there is no confirmation prompt, logging, or explicit warning in the docstring beyond a neutral description. Because this is a code file, safety-relevant filesystem modification should have some visible disclosure unless clearly warned elsewhere.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.