Back to skill

Security audit

Invoice Scan

Security checks for vulnerabilities and agentic risk

Overview

The skill matches invoice scanning, but it needs review because it handles sensitive invoices and has concrete security flaws in retry merging, CSV export, and remote upload scoping.

Review before installing or using on confidential invoices. CLI mode sends full invoice documents to an external AI provider, so use only approved documents and credentials. Avoid exposing this package as a long-running service until retry path validation, CSV formula neutralization, input size limits, endpoint restrictions, and dependency pinning are addressed.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/validation/completeness.js:229
Finding

Model-Controlled Property Path Enables Prototype Pollution

Content
View full analysis
Remediation
View remediation
{ invoice.header.invoiceNumber = value; }, 'header.invoiceDate': (invoice, value) => { invoice.header.invoiceDate = value; }, 'totals.grossTotal': (invoice, value) => { invoice.totals.grossTotal = value; }, }; function mergeRetryResults(invoice, retryResults, requestedPaths) { const allowed = new Set(requestedPaths); for (const result of retryResults) { if (!allowed.has(result.path)) { continue; } const setter = RETRY_SETTERS[result.path]; if (!setter || result.value == null || result.value === NOT_ON_DOCUMENT) { continue; } setter(invoice, result.value); } } ``` 4. Validate retry values against the expected type and format for each field before assignment. 5. Add tests using paths such as: - `__proto__.polluted` - `constructor.prototype.polluted` - `header.__proto__.polluted` - Valid but unrequested schema paths 6. Treat all model output as attacker-influenced input, even when the prompt requests an exact JSON structure. ]]>

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/output/csv.js:11
Finding

CSV Formula Injection Through Untrusted Invoice Fields

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Note
Location
scripts/extraction/scanner.js:51
Finding

Unbounded Document Loading and Image Decoding Enables Resource Exhaustion

Content
View full analysis
Remediation
View remediation
MAX_FILE_BYTES) { throw new Error(`Input file exceeds the ${MAX_FILE_BYTES}-byte limit`); } const imageBuffer = await fs.promises.readFile(filePath); const pipeline = sharp(imageBuffer, { limitInputPixels: MAX_PIXELS, sequentialRead: true, }); ``` 5. Validate the actual file signature before selecting an image or PDF processing path. 6. Reject unsupported multi-page or unusually large PDF files before transmission. 7. Add request-level timeouts and concurrency controls when the scanner is exposed through a server. 8. Test oversized files, extreme image dimensions, malformed images, compressed TIFF files, and concurrent processing workloads. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (14)

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/cli.js (reported line 49)May include surrounding context.

js
const noPreprocess = hasFlag('no-preprocess');
      const acceptTypes = getArg('accept', 'relaxed'); // strict, relaxed, any

      // Get API key from args or env
      let apiKey = getArg('api-key');
      if (!apiKey) {
        const envMap = {

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · scripts/test/run-tests.js (reported line 227)May include surrounding context.

js
assert(arith7.validation.arithmeticValid === true, 'amountDue check: arithmetic still valid (warning not error)');
assert(arith7.validation.warnings.some(w => w.field === 'totals.amountDue'), 'amountDue mismatch: warning generated');

// 2h. amountDue correct — no warning
const arith8 = makeFullInvoice({ totals: { amountPaid: -80, amountDue: 100, grossTotal: 180 } });
validateArithmetic(arith8);
assert(!arith8.validation.warnings.some(w => w.field === 'totals.amountDue'), 'amountDue correct: no warning');

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · scripts/test/run-tests.js (reported line 230)May include surrounding context.

js
assert(arith7.validation.arithmeticValid === true, 'amountDue check: arithmetic still valid (warning not error)');
assert(arith7.validation.warnings.some(w => w.field === 'totals.amountDue'), 'amountDue mismatch: warning generated');

// 2h. amountDue correct — no warning
const arith8 = makeFullInvoice({ totals: { amountPaid: -80, amountDue: 100, grossTotal: 180 } });
validateArithmetic(arith8);
assert(!arith8.validation.warnings.some(w => w.field === 'totals.amountDue'), 'amountDue correct: no warning');

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill declares capabilities that involve environment-variable access and external network transmission, but it does not define an explicit tool scope such as permissions or allowed-tools. That creates an authorization ambiguity: a host agent may permit broader execution or fail to enforce least-privilege controls, especially since the skill can transmit sensitive invoice data and use an API key in CLI mode.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The CLI processes local invoice files and sends their contents to an external AI provider, but the user-facing flow shown here does not provide an explicit warning or consent checkpoint before transmission. Because invoices commonly contain sensitive financial and personal data, silent off-device transfer can create privacy, confidentiality, and compliance risk even if the transfer is part of intended functionality.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The scanner sends invoice contents to an external AI provider as part of its normal OCR/extraction workflow, but this file contains no explicit disclosure, consent gate, or locality control before transmitting potentially sensitive financial documents. Because invoices commonly contain PII, payment details, addresses, tax IDs, and bank information, silent third-party transmission creates a real privacy and compliance risk even if the behavior is intentional product functionality.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

This code base64-encodes the full document and POSTs it to a configurable remote endpoint via apiBaseUrl, which means complete invoice contents are exfiltrated off-host for processing. The risk is heightened because apiBaseUrl is user-configurable with no visible restriction or validation here, enabling transmission of sensitive documents to arbitrary endpoints if misused or misconfigured.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

This code generates an XLSX buffer containing invoice metadata, addresses, VAT IDs, and bank details such as IBAN, BIC, account number, and sort code. Although the file has technical comments, it does not include any user-facing warning, confirmation, or disclosure that sensitive financial data will be included in the exported workbook.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This manifest file says the skill can "extract structured data from any invoice document," but it does not provide any specific activation phrases, invocation constraints, or exclusion conditions. In a manifest context, that broad natural-language scope can make it unclear when the skill should activate versus when general document-related conversation should not invoke it.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
90% confidence
Finding

Using a caret range for sharp allows future installs to resolve to different releases, which weakens reproducibility and can silently introduce vulnerable or behavior-changing versions into the supply chain. In an OCR/document-processing skill that handles untrusted files, dependency drift is more dangerous because image parsing libraries are high-risk attack surfaces.

Content

Scanner excerpt · scripts/package.json (reported line 37)May include surrounding context.

json
"node": ">=18.0.0"
  },
  "dependencies": {
    "sharp": "^0.33.0",
    "xlsx": "^0.18.5"
  }
}

Unverifiable Dependency: sharp has 4 known advisory(ies) (GHSA-54xq-cgqr-rpm3 (sharp vulnerability in libwebp dependency CVE-2023-4863); GHSA-f88m-g3jw-g9cj (sharp inherited vulnerabilities in libvips: CVE-2026-33327, CVE-2026-33328, CVE-); CVE-2022-29256 (sharp vulnerable to Command Injection in post-installation over build environmen) +1 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
88% confidence
Finding

The manifest does not pin sharp, while the dependency family has known advisories affecting image-processing components and prior install-time issues. Because this skill processes untrusted invoice images and PDFs, any vulnerable image codec or native library path could expose denial-of-service, memory corruption, or install-time compromise risk.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
93% confidence
Finding

Using a caret range for xlsx permits non-deterministic installs and increases supply-chain risk, including accidental adoption of vulnerable parser behavior. This is especially relevant here because spreadsheet/document parsing processes attacker-controlled invoice content, making parser bugs such as prototype pollution or ReDoS more impactful.

Content

Scanner excerpt · scripts/package.json (reported line 38)May include surrounding context.

json
},
  "dependencies": {
    "sharp": "^0.33.0",
    "xlsx": "^0.18.5"
  }
}

Unverifiable Dependency: xlsx has 5 known advisory(ies) (CVE-2021-32012 (Denial of Service in SheetJS Pro); CVE-2023-30533 (Prototype Pollution in sheetJS); CVE-2024-22363 (SheetJS Regular Expression Denial of Service (ReDoS)) +2 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
95% confidence
Finding

The manifest leaves xlsx unpinned despite multiple known advisories in that ecosystem, and the package is used to parse or generate spreadsheet-like structured output from attacker-influenced document content. If an affected version is installed, issues such as prototype pollution or ReDoS could lead to application compromise or service disruption during document processing.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
98% confidence
Finding

The docstring for preprocessImage says it returns a processed image buffer in PNG format, but the code later forces JPEG output with pipeline.jpeg({ quality: 90 }). This is an active contradiction in the file's own documentation about the function's behavior and output format.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.