Back to skill

Security audit

Invoice Extractor

Security checks for vulnerabilities and agentic risk

Overview

This looks like a legitimate invoice and expense-ledger skill, but it deserves review because it can persistently change financial records and save export settings without enough confirmation or safety checks.

Install only if you are comfortable letting the skill read invoice files and maintain a local CSV expense ledger. Review extracted entries before adding them, keep backups, be careful with delete/undo/edit commands, avoid saving web-discovered export presets without checking them, and treat generated CSV files as untrusted when opening them in spreadsheet software.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:31
Finding

Unpinned Third-Party Dependencies Allow Supply-Chain Substitution

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:31; related installation guidance in references/notes.md:23-24 and scripts/extract.py:135-136
Vulnerability Type: Unpinned Python package installation
Risk Level: Medium

Vulnerable Code

bash
pip install pdfplumber

Related fallback instructions:

markdown
- **pdfplumber** — primary PDF extraction: `pip install pdfplumber`
- **PyPDF2** — fallback if pdfplumber unavailable: `pip install PyPDF2`

The executable also prints the same unpinned installation commands:

python
print("  pip install pdfplumber", file=sys.stderr)
print("  # or: pip install PyPDF2", file=sys.stderr)

Technical Analysis

The installation instructions request packages by name without specifying an audited version or validating package hashes. Consequently, the code installed by a user depends on whichever release the package index serves at installation time rather than the release reviewed with this project.

This is a supply-chain weakness rather than evidence that the current pdfplumber or PyPDF2 packages are malicious. Exploitation would require compromise of an upstream package, maintainer account, package-distribution channel, or resolution environment. A compromised distribution can execute code during installation through build hooks or later when imported by scripts/extract.py.

Attack Path

  1. An attacker compromises an upstream dependency release, maintainer account, package index, or package-resolution environment.
  2. The victim follows the documented pip install pdfplumber or fallback pip install PyPDF2 instruction.
  3. Because no version or cryptographic hash is enforced, pip retrieves the attacker-controlled release.
  4. Malicious code executes during package build or installation, or when extract_pdf_text() imports the package.
  5. The payload operates with the permissions of the user or service running pip and the in ...[truncated 510 chars]
Remediation
View remediation

Remediation Suggestions

  1. Pin each supported dependency to a reviewed version, for example through a committed requirements file:
    text
    pdfplumber==<reviewed-version>
    PyPDF2==<reviewed-version>
    
  2. Generate and enforce cryptographic hashes:
    bash
    python3 -m pip install --require-hashes -r requirements.txt
    
  3. Pin transitive dependencies using a lock-file workflow such as pip-tools, Poetry, or uv.
  4. Replace every installation example and runtime error message with the locked installation command.
  5. Run dependency vulnerability and provenance checks in CI, and review automated dependency updates before merging.
  6. Install into an isolated virtual environment under an unprivileged account; do not recommend sudo pip.
  7. Prefer a trusted, explicitly configured package index and retain verified dependency artifacts where reproducible deployment is required.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/extract.py:714
Finding

Attacker-Controlled Invoice Fields Are Exported Without Spreadsheet Formula Neutralization

Content
View full analysis

Vulnerability Details

File Location: scripts/extract.py:714-775, with CSV output sinks at scripts/extract.py:564-569 and scripts/extract.py:827-833
Vulnerability Type: CSV/spreadsheet formula injection
Risk Level: Medium

Vulnerable Code

Invoice-controlled ledger values are copied directly into export rows:

python
if transform == "xero":
    account_code = config.get("xero", {}).get("defaultAccountCode", "200")
    tax_rate = config.get("xero", {}).get("defaultTaxRate", "23% (VAT on Expenses)")
    total = float(row.get("total", 0) or 0)
    return [
        row.get("vendor", ""),
        row.get("id", ""),
        format_date(row.get("date", ""), date_fmt),
        format_date(row.get("dueDate", ""), date_fmt) or format_date(row.get("date", ""), date_fmt),
        row.get("description", ""),
        "1",
        f"{total:.2f}",
        account_code,
        tax_rate,
    ]

Custom presets likewise return untrusted values unchanged:

python
elif transform == "custom":
    field_mapping = preset.get("fieldMapping", {})
    result = []
    for col in preset.get("columns", []):
        ledger_field = field_mapping.get(col, col)
        val = row.get(ledger_field, "")
        if ledger_field == "date" and val:
            val = format_date(val, date_fmt)
        if ledger_field == "total" and val:
            total = float(val)
            if amount_handling == "negative":
                val = f"-{total:.2f}"
            else:
                val = f"{total:.2f}"
        result.append(str(val))
    return result

The transformed values are then written directly to CSV:

python
if preset.get("headerRow", True):
    writer.writerow(preset.get("columns", []))

for row in filtered:
    writer.writerow(transform_row(row, preset, config))

The CSV view path similarly emits raw ledger rows:

python
writer = csv.DictWriter
...[truncated 2337 chars]
Remediation
View remediation

Remediation Suggestions

  1. Apply a centralized CSV-safety function to every untrusted textual cell before all CSV output:
    python
    FORMULA_PREFIXES = ("=", "+", "-", "@", "\t", "\r", "\n")
    
    def sanitize_csv_cell(value):
        text = str(value)
        if text.startswith(FORMULA_PREFIXES):
            return "'" + text
        return text
    
  2. Use the sanitizer in transform_row(), custom field mappings, ledger_view(..., fmt="csv"), and preset column headers if custom configuration is not fully trusted.
  3. Preserve numeric types through explicit validation rather than sanitizing legitimate numeric values as text. Reject non-finite amounts such as NaN and infinity.
  4. Consider rejecting dangerous leading characters during ledger ingestion for fields that never legitimately require formulas.
  5. Normalize or reject leading control characters before testing the first meaningful character.
  6. Document that generated files contain inert data and test exports against Excel, LibreOffice Calc, and Google Sheets import behavior.
  7. Add regression tests covering =, +, -, @, tabs, carriage returns, newlines, quoted values, and every built-in and custom export preset.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (18)

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The documented behavior is broader and materially different from the declared extraction purpose: the agent is instructed to perform parsing itself, manage a persistent ledger, mutate entries, delete data, export to external formats, and rely on other tools for image handling. This mismatch can cause unsafe invocation in contexts where a user expects passive extraction but the skill can also modify stored financial data or initiate additional workflows.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill directs the agent to read invoice files and write ledger/config/export files, but it declares no explicit tool scope or allowed-tools boundaries. That increases the chance the skill is invoked with broader file access than intended, making unintended reads or persistent writes harder to constrain or audit.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest uses broad finance-related activation language that can match generic requests like expense tracking or tax prep without tightly limiting scope. Over-broad triggering can cause the skill to activate in situations where file access, ledger mutation, or exports were not the user's intent, increasing the chance of unintended sensitive-data handling.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The 'When to Use' examples include vague phrases like 'Track my expenses' and 'Show me my spending summary,' which are broad enough to capture many finance requests. In this skill context that matters because the skill includes persistent ledger actions and data exports, so loose routing can expose sensitive financial data or trigger unintended stateful operations.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill instructs the agent to use web search to discover third-party CSV import formats and then operationalize that information. This expands the trust boundary from local invoice processing to external, potentially unreliable content, creating a path for prompt injection, bad mappings, or unsafe config changes based on unverified web results.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill tells the agent to persist a newly discovered export preset into the user's config file without an explicit confirmation step. Any persistent config write changes future behavior and could silently corrupt exports, embed incorrect mappings, or create a durable foothold for bad settings derived from untrusted web content.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

L032 states that the skill requires no external APIs, but L037 says the agent will search platform documentation for unknown export targets and persist new presets. Searching external docs implies network access or an equivalent external-retrieval capability, which exceeds the stated no-external-APIs behavior.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The architecture section documents dependencies as 'pdfplumber (pip), stdlib for everything else,' suggesting a local/offline implementation. Elsewhere, the plan includes LLM-powered doc discovery (L037) and future exchange-rate/API integrations (L166, L170-L173), which contradict the stated implementation boundaries rather than merely omitting details.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This markdown file documents commands that remove or rewrite ledger data, including deletion, editing with hash recalculation, and undoing the last entry. Under the markdown-file criteria for missing user warnings, these destructive behaviors should be accompanied by an explicit warning about data modification or reversibility risks.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The roadmap proposes direct API pushes to accounting systems plus credential storage, but provides no security constraints around token protection, consent, scope limitation, or data handling. In a financial-document skill, this creates meaningful risk of unauthorized account actions, credential exposure, and leakage of sensitive accounting data if implemented carelessly.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Automated Gmail scanning and invoice processing implies broad access to inbox contents and attachment data, which may include sensitive financial and personal information. Without explicit warnings, consent boundaries, mailbox scope restrictions, and automation safeguards, users could unknowingly grant extensive data access or trigger unintended processing of private documents.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The delete command removes a selected ledger entry and rewrites the ledger file, which is a destructive file-modification operation. Although the script prints what was deleted afterward, there is no prior confirmation prompt or explicit warning before the irreversible change is made.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The undo command deletes the highest-ID ledger entry and rewrites the ledger CSV, affecting user data in a destructive way. The function only reports the removal after it has already happened, so users are not warned or asked to confirm before the change.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
80% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · scripts/extract.py (reported line 872)May include surrounding context.

python
ledger_sub = ledger_parser.add_subparsers(dest="ledger_command")

    # ledger add
    add_parser = ledger_sub.add_parser("add", help="Add entry to ledger")
    add_parser.add_argument("json_file", nargs="?", help="JSON file path (or - for stdin)")
    add_parser.add_argument("--source", help="Source filename for the entry")
    add_parser.add_argument("--force", action="store_true", help="Skip duplicate detection")

Vague Triggers

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

This manifest-style JSON defines many generic keywords such as "hotel," "taxi," "equipment," "subscription," and "restaurant" without any documented scope limits or exclusion conditions. In a manifest/config context, such broad terms can create ambiguous matching behavior and unintended invocations or classifications.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The defaults set currency to EUR and date format to DD/MM/YYYY, which imposes a specific regional convention. Under the policy, locale-specific behavior should either be user-selectable or clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

Encouraging receipt submission over third-party messaging platforms can expose financial documents, metadata, and account-linked personal information to additional processors and insecure user workflows. In the invoice-processing context, the absence of a privacy warning increases the chance that users share sensitive documents through channels with retention, forwarding, or device-compromise risks.

Content

No source excerpt is available for this finding.

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · scripts/extract.py (reported line 937)May include surrounding context.

python
updates = {}
            for field in ("vendor", "total", "date", "description", "category",
                          "currency", "subtotal", "tax"):
                val = getattr(args, field, None)
                if val is not None:
                    updates[field] = val
            if not updates:

Static analysis

No suspicious patterns detected.