Back to skill

Security audit

Invoice-Recognition

Security checks for vulnerabilities and agentic risk

Overview

The skill appears to perform invoice OCR as advertised, but it needs review because it mishandles API credentials and can create unsafe Excel files from untrusted invoice text.

Install only after reviewing the credential handling and Excel-output risks. Do not use the sample keys in setup.md, prefer environment variables or a secret manager over config.txt or command-line secrets, confirm you are allowed to send the invoices to Baidu OCR, and treat generated XLSX reports as sensitive files that may need formula sanitization before opening or sharing.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
src/excel_exporter.py:62
Finding

Spreadsheet Formula Injection Through OCR-Controlled Invoice Fields

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
setup.md:168
Finding

Credential-Shaped Secrets in Documentation and Unsafe Credential Handling

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
src/invoice_extractor.py:117
Finding

Predictable Temporary Files Allow Symlink Overwrite and Cross-Process Collisions

Content
View full analysis
0.5: all_text_lines.append(text) # Clean temporary file temp_img_path.unlink(missing_ok=True) ``` The image-to-PDF fallback also derives a predictable filename from attacker-influenced input: ```python # Save temporary PDF file temp_pdf_path = Path(".temp/cache") / f"{image_path.stem}.pdf" temp_pdf_path.parent.mkdir(parents=True, exist_ok=True) with open(temp_pdf_path, 'wb') as f: f.write(pdf_bytes.getvalue()) # Use PDF extraction method invoice = self._extract_from_pdf(temp_pdf_path) # Update source file path if invoice: invoice.source_file = str(image_path) # Clean temporary file temp_pdf_path.unlink(missing_ok=True) ``` ### Technical Analysis The code creates temporary files in a predictable relative directory and opens them with normal write mode. Normal `open(..., "wb")` follows symbolic links and truncates existing files. The PDF-page name depends only on the page number, while the fallback PDF name depends on the source filename stem. As a result: - Two simultaneous runs can overwrite or delete each other’s files. - Multiple files processed concurrently can collide. - A local attacker who can prepare the working directory can create a symbolic link at an expected path. - Cleanup can unlink an attacker-selected directory entry ...[truncated 1603 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Installer Uses Unbounded Dependency Versions Without Integrity Verification

Content
View full analysis
=2.28.0 pandas>=2.0.0 openpyxl>=3.1.0 PyMuPDF>=1.23.0 Pillow>=10.0.0 ``` The installer resolves and installs those mutable ranges: ```bash # Install dependencies echo "" echo "Installing dependencies..." pip install -r requirements.txt ``` ### Technical Analysis Every dependency specifies only a minimum version. A fresh installation can therefore select any future release satisfying the constraint. The reviewed source tree does not determine the exact code that will be installed and executed. Python package installation can run package build logic, and the installed libraries are later imported by the Skill. A compromised upstream release, account takeover, malicious mirror, or newly incompatible release can consequently alter the effective behavior after this audit. The installer also does not: - Use a lock file. - Require package hashes. - Select a controlled package index. - Create or require an isolated virtual environment. - Stop execution explicitly if dependency installation fails. No typosquatted package name or known malicious package was identified in the reviewed requirements; the risk arises from mutable, unverified resolution. ### Attack Path 1. A future compromised or malicious release is published under one of the permitted package names and satisfies its `>=` constraint. 2. A user runs `install.sh`. 3. `pip` resolves the new release from the configured package index or mirror. 4. Package build or installation logic executes with the invoking user’s permissions. 5. The malicious package is imported during normal Skill execution. 6. The dependency can access invoice files, API credentials, generated reports, and any other resources accessible to the process. A malicious package index or compromised mirror can ...[truncated 777 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (33)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Claiming Baidu OCR while using a different OCR engine, and claiming Excel export and multi-file support without clearly implementing them, can cause users to expose sensitive invoice data under false assumptions about where data goes and what processing occurs. Such mismatches are especially risky for financial documents because they affect privacy expectations, compliance obligations, and trust in downstream outputs.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

Claiming Baidu OCR while using a different OCR engine, and claiming Excel export and multi-file support without clearly implementing them, can cause users to expose sensitive invoice data under false assumptions about where data goes and what processing occurs. Such mismatches are especially risky for financial documents because they affect privacy expectations, compliance obligations, and trust in downstream outputs.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill processes invoices and receipts, which commonly contain sensitive financial and personal data, yet it does not clearly warn users that document contents will be transmitted to Baidu OCR. Without explicit disclosure, users cannot make an informed decision about sharing regulated or confidential data with a third party, creating privacy, compliance, and data-handling risk.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The setup guide embeds real-looking Baidu API and Secret Key values directly in a manual configuration example, and they are not labeled as fake placeholders. Readers may copy and reuse them, which can expose a real third-party account to unauthorized API consumption, billing abuse, service suspension, or compromise if the credentials are valid.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

This is a true secret-exposure issue in documentation: the example shows realistic credential strings without any warning that they are nonfunctional placeholders. In the context of a skill specifically designed to interact with Baidu OCR, such keys are especially dangerous because users are expected to use them for authentication, making accidental misuse or exploitation more likely.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
78% confidence
Finding

The skill advertises capabilities that inherently require file access, environment access, and network transmission to a third-party OCR provider, but it does not declare an explicit tool scope or permissions boundary. This creates a transparency and governance gap: users and the agent framework cannot clearly constrain or audit what resources the skill may access, increasing the risk of overbroad file reads, credential use, and external data transfer.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

An overly broad trigger can cause the skill to activate for general invoice or receipt requests where users did not intend OCR processing or third-party transmission. In this context, mistaken activation is more dangerous because invoice and receipt files often contain financial identifiers, addresses, tax numbers, and other sensitive business data.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The examples encourage uploading invoice PDFs/images to Baidu OCR but provide no user-facing notice that invoice contents will be transmitted to a third-party service. Invoices routinely contain sensitive financial and personal data, so omission of a privacy/data-sharing warning can lead to unintended disclosure and compliance issues.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

Passing API keys and secret keys directly on the command line can expose credentials through shell history, process listings, audit logs, and CI job logs. This creates a realistic path for credential theft and misuse of the OCR account.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The automated email example sends generated Excel invoice reports as attachments without warning that these files may contain sensitive financial and personal information. Users may adopt the pattern directly and exfiltrate data to the wrong recipient, insecure mailbox, or unmanaged email environment.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The module docstring presents the skill entirely in Chinese, and the command-line interface strings throughout the file are also Chinese-only. This imposes a specific language/locale on users without offering an opt-in choice or documenting a justified region-specific constraint.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The module docstring and all user-facing CLI messages are written exclusively in Chinese, indicating a fixed language experience. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation unless the locale restriction is explicitly documented and justified, which it is not here.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill sends invoice images/PDF-derived image data to Baidu's external OCR API, which means potentially sensitive financial and personal data leaves the local environment. In an invoice-processing skill, this is expected functionality, but the lack of an explicit user-facing disclosure/consent mechanism creates a real privacy and compliance risk because users may not realize documents are being uploaded to a third party.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Comments, docstrings, prompts, and status messages are consistently written in Chinese, including the interactive setup flow. This creates a language policy concern because the skill does not offer any language choice or indicate that it is intentionally limited to Chinese users.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The code persists BAIDU_API_KEY and BAIDU_SECRET_KEY directly into config.txt in plaintext. If the working directory is shared, backed up, committed to source control, or readable by other local users/processes, these credentials can be exposed and then abused to access the Baidu OCR account or consume its quota.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The module description and all user-facing strings are in Chinese, including status messages and default worksheet/file names. This imposes a specific language/locale on users without offering a language choice or documenting that the skill is intended only for a Chinese-speaking or region-specific context.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The code explicitly initializes PaddleOCR with lang='ch', and the surrounding natural-language descriptions are also Chinese-only. Under the policy, forcing a specific language or locale without user opt-in is a natural-language policy concern unless the constraint is clearly documented and justified as region-specific.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

When OCR fallback is used for PDF pages, rendered page images containing invoice data are written to .temp/cache on local disk. Invoice documents commonly contain sensitive financial and personal/business information, so persisting them without clear user notice or secure handling increases the risk of unintended disclosure through leftover files, backups, shared workspaces, or permissive filesystem access.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The image fallback path converts the source image into a temporary PDF and stores it under .temp/cache, again persisting invoice contents locally. Because this skill is specifically designed for invoice extraction, the temporary file is very likely to contain regulated or confidential billing data, making accidental retention or exposure materially risky.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The module docstring and CLI-facing description are written entirely in Chinese, which imposes a specific language on users without any opt-in or explanation. Under the policy rule, forcing a language/locale without user choice or documented justification is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This tool processes invoices, which commonly contain sensitive financial, tax, and identity information, and sends them to a third-party OCR provider while also storing extracted results locally in Excel. The file does not present a clear privacy notice, consent step, or data-handling warning, so users may unknowingly expose regulated or confidential data to external services and local storage.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
98% confidence
Finding

The dependency is specified with a lower bound only, so installs are not reproducible and may pull in unexpected future versions. In a skill that processes untrusted invoice images/PDFs and makes network requests, dependency drift can silently introduce vulnerable or breaking releases into the runtime.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
requests>=2.28.0
pandas>=2.0.0
openpyxl>=3.1.0
PyMuPDF>=1.23.0

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
95% confidence
Finding

Because requests is not pinned, it is impossible to verify from this manifest whether the deployed version includes fixes for known advisories. The risk is somewhat contextual: this skill likely calls an external OCR API, so an unsafe HTTP client version could affect credential handling, redirects, or TLS-related behavior depending on the actual installed release.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
97% confidence
Finding

Using an unpinned pandas version means the installed package can vary across environments and over time, reducing reproducibility and making it hard to verify exposure to known issues. Although this file alone does not prove exploitation, it increases supply-chain and maintenance risk.

Content

Scanner excerpt · requirements.txt (reported line 2)May include surrounding context.

text
requests>=2.28.0
pandas>=2.0.0
openpyxl>=3.1.0
PyMuPDF>=1.23.0
Pillow>=10.0.0

Unverifiable Dependency: pandas has 1 known advisory(ies) (CVE-2020-13091 (** DISPUTED ** pandas through 1.0.3 can unserialize and execute commands from an)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
90% confidence
Finding

The manifest does not allow verification that the installed pandas release is free of known issues. Even though the cited pandas advisory may be situational, leaving the version unpinned prevents reliable assessment and can expose the skill to vulnerable dependency resolution.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
examples.md:315

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/batch_process.py:88

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
setup.md:170