Back to skill

Security audit

发票内容识别

Security checks for vulnerabilities and agentic risk

Overview

This appears to be a real Baidu-based invoice OCR skill, but it needs review because it handles sensitive invoices and API credentials while using unsafe installation and credential-handling practices.

Install only if you are comfortable sending invoice images and extracted tax/business data to Baidu Cloud. Use a dedicated virtual environment, pin dependencies, avoid system Python modification, keep Baidu credentials separate and rotated, and do not process invoices that cannot legally or contractually leave your environment.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:40
Finding

Unpinned dependencies installed outside an isolated environment

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:40
Vulnerability Type: Unsafe third-party dependency installation
Risk Level: Medium

Vulnerable Code

bash
pip install pymupdf openpyxl requests Pillow python-dotenv --break-system-packages -q

Technical Analysis

The installation command retrieves dependencies without pinning reviewed versions or verifying package hashes. Consequently, the code installed during each invocation may differ from the versions originally reviewed.

The --break-system-packages option bypasses protections intended to prevent pip from modifying a system-managed Python environment. This exceeds the minimum privileges necessary for the Skill because its dependencies can instead be installed in a dedicated virtual environment. The -q option also suppresses installation details that could help users identify unexpected package sources or dependency changes.

No dependency-confusion package name or known malicious dependency was identified in the reviewed project. The risk arises from unsafe supply-chain and environment-management practices rather than evidence that the listed packages are currently malicious.

Attack Path

  1. An attacker compromises a listed package, one of its transitive dependencies, or the package distribution channel.
  2. A user follows the installation command in SKILL.md.
  3. Because versions and hashes are not constrained, pip resolves and downloads the attacker-controlled release.
  4. Package installation or subsequent import executes the malicious component with the permissions of the user running the command.
  5. Because installation is allowed to modify the system-managed Python environment, the compromised component may affect other Python applications using that environment.

Impact Assessment

Successful exploitation could execute code with the invoking user's privileges, access files and credentials available to that user, alter ...[truncated 298 chars]

Remediation
View remediation

Remediation Suggestions

  • Create and use a dedicated virtual environment rather than passing --break-system-packages.
  • Pin direct and transitive dependencies to reviewed versions in a lock file.
  • Generate and enforce cryptographic hashes, for example with a hash-locked requirements file and pip install --require-hashes.
  • Configure an explicit trusted package index and review transitive dependencies.
  • Remove -q so package resolution, source, and installation failures remain visible.
  • Run dependency installation and invoice processing as a non-privileged user.
  • Add automated dependency vulnerability and provenance scanning to the release process.

A hardened workflow should resemble:

bash
python -m venv .venv
. .venv/bin/activate
python -m pip install --require-hashes -r requirements.lock

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/invoice_ocr_main.py:228
Finding

API credentials and bearer tokens are transmitted in URL query strings

Content
View full analysis

Vulnerability Details

File Location: scripts/invoice_ocr_main.py:228-229, 270-278
Vulnerability Type: Sensitive authentication data in request URLs
Risk Level: Medium

Vulnerable Code

python
params = {"grant_type": "client_credentials", "client_id": self.api_key, "client_secret": self.secret_key}
resp = requests.post(url, params=params, timeout=15)
python
vat_url = f"https://aip.baidubce.com/rest/2.0/ocr/v1/vat_invoice?access_token={token}"
resp = requests.post(vat_url, headers=headers, data={"image": img_b64}, timeout=30)
resp.raise_for_status()
result = resp.json()

if result.get("error_code") and result["error_code"] != 0:
    print("    ⚠️  增值税专用接口失败,降级通用OCR...")
    gen_url = f"https://aip.baidubce.com/rest/2.0/ocr/v1/general_basic?access_token={token}"
    resp2 = requests.post(gen_url, headers=headers, data={"image": img_b64}, timeout=30)

Technical Analysis

Passing params to requests.post places client_id and client_secret in the token endpoint's URL query string. The OCR endpoint URLs similarly interpolate the bearer access token directly into the query string.

HTTPS encrypts the request in transit and the reviewed code does not explicitly print these URLs. Nevertheless, URLs are commonly captured by HTTP diagnostics, reverse proxies, endpoint monitoring, exception-reporting systems, or server access logs. Query-string secrets therefore have a larger accidental-disclosure surface than credentials carried in protected authorization headers or request bodies.

The .env file included in the project contains placeholders rather than functional credentials, so no bundled live secret was identified. The affected credentials would be supplied by the user at runtime.

The same function Base64-encodes complete invoice images and uploads them to Baidu Cloud OCR:

python
img_b64 = base64.b64encode(self._prepare_for_baidu(image_bytes)).decode("utf-8")

...[truncated 1772 chars]

Remediation
View remediation

Remediation Suggestions

  • Follow the current Baidu API authentication specification and use an authorization header or POST body whenever the provider supports it.
  • If the provider mandates query-string credentials, isolate the request code and ensure URLs are never included in application logs, proxy logs, traces, exceptions, or monitoring events.
  • Add a redaction filter for client_secret, access_token, client_id, and authorization values before recording HTTP metadata.
  • Use narrowly scoped credentials where supported and rotate the client secret regularly.
  • Keep access tokens only in memory, minimize their lifetime, and clear cached tokens when processing ends.
  • Disable verbose HTTP-library debugging in production.
  • Display an explicit notice and obtain user consent before uploading invoice images to the third-party OCR service.
  • Document the destination, transmitted fields, retention policy, and relevant data-processing terms.
  • Consider a local OCR option for invoices that cannot legally or contractually be transferred to an external service.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (11)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

整体上,这段代码的主用途与“增值税发票 OCR 识别并导出 Excel”这一大方向是一致的,因此不是完全无关的技能。但声明对能力描述明显强于实际实现,存在重要行为差异。代码支持 PDF/图片输入、百度 OCR 调用、结果导出,这些核心方向是匹配的;不过其“发票预检测”只是简单尺寸校验,真正判断依赖 OCR 结果后的关键词匹配,难称为完整预检测。质量评估模块仅使用必填字段覆盖率和检测置信度计算总分,数据结构中虽保留 format_score、amount_consistency_score 等字段,但并未真正实现相应逻辑。Excel 导出也只有一个工作表“发票汇总”,没有声明中所述更完整的多维报告输出。另一个差异是字段提取只覆盖发票类型、号码、日期、买卖方和金额等少数字段,未达到声明给人的完整结构化提取预期。因此应判定为描述与实际能力存在实质性不匹配,属于夸大实现范围而非完全错误描述。

Content

No source excerpt is available for this finding.

Vague Triggers

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The trigger phrases are overly broad, including wording that says even a simple request like 'help me recognize this invoice' should always activate the skill. Over-broad activation can cause unintended processing of sensitive financial documents and unnecessary transmission of data to a third-party OCR provider without clear user intent or informed consent.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
90% confidence
Finding

The skill explicitly instructs reading credentials from .env, obtaining and caching an access token, and using it for external API calls. Credential access is security-sensitive; without strict scoping, storage, and handling controls, tokens or API secrets may be exposed through logs, wider environment access, or misuse by adjacent components.

Content

Scanner excerpt · SKILL.md (reported line 71)May include surrounding context.

md
### 3. 内容提取 (ContentExtractor)

* 严格通过 `python-dotenv` 读取鉴权信息,获取并缓存 Access Token。
* 优先请求 `vat_invoice` 增值税专用接口。
* 若接口返回错误码(如非标准发票类型),无缝降级调用 `general_basic` 通用文字识别接口,并通过正则尽力挽救关键数据。
* 内置自动重试机制(最高 3 次,间隔递增),增强网络抗抖动能力。

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/invoice_ocr_main.py (reported line 27)May include surrounding context.

python
from PIL import Image
from dotenv import load_dotenv

# 加载 .env 配置文件
load_dotenv()

# ─────────────────────────────────────────────────────────────────────────────

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/invoice_ocr_main.py (reported line 218)May include surrounding context.

python
from PIL import Image
from dotenv import load_dotenv

# 加载 .env 配置文件
load_dotenv()

# ─────────────────────────────────────────────────────────────────────────────

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill performs sensitive capabilities—reading environment variables, accessing local files, and making network requests—but does not declare any explicit tool scope or permissions boundary. This creates ambiguity for operators and increases the risk of over-privileged execution, especially because the skill instructs reading .env credentials and transmitting document contents to an external OCR service.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The skill is described exclusively as a Chinese VAT invoice recognition workflow centered on Baidu's VAT invoice API, with activation keywords and outputs oriented to Chinese-language usage, but it does not offer a language or locale choice. This creates a natural-language locale constraint without explicit user opt-in or a clearly documented regional limitation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill description does not clearly warn users that invoice images/PDFs and extracted financial data will be sent to Baidu's OCR service. Because invoices often contain sensitive personal, tax, and business information, lack of disclosure undermines consent and can lead to privacy, compliance, and data-handling violations.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script sends invoice images, which typically contain sensitive financial and tax information, to Baidu Cloud OCR services without any explicit user notice, consent flow, or privacy guardrails. In an agent skill context, users may assume local processing, so silent third-party transmission can cause confidentiality, compliance, and data-handling violations.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The top-level description states '仅支持百度云OCR' in a Chinese-language workflow for VAT invoices, and the script presents its interface and behavior entirely in Chinese without offering a language or locale option. Under the policy, forcing a specific language or locale without user opt-in can be a natural-language policy violation unless clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · scripts/invoice_ocr_main.py (reported line 338)May include surrounding context.

python
result.grade = "🔴 失败/非发票"
            return result

        req_present = sum(1 for f in self.REQUIRED_FIELDS if getattr(content, f, ""))
        result.completeness_score = (req_present / len(self.REQUIRED_FIELDS)) * 100
        result.clarity_score = detection.confidence * 100

Static analysis

No suspicious patterns detected.