Back to skill

Security audit

Invoice Guard

Security checks for vulnerabilities and agentic risk

Overview

This invoice skill is mostly purpose-aligned, but it can produce misleading compliance reports that treat unverified invoices as normal while handling sensitive financial data.

Review this skill before installing. Do not rely on its reports as proof that invoices were officially verified unless a real tax-platform integration is added and unchecked records are clearly labeled. Use explicit consent before sending invoice data to Feishu or tax services, and treat generated Markdown or Feishu reports as sensitive business records.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/compliance_report.py:281
Finding

Unverified invoices are falsely reported as verified and normal

Content
View full analysis
tuple[str, str]: """ Call the tax verification platform to verify invoice authenticity. Returns: (verification status, message) """ allowed, msg = tier.allow_verify() if not allowed: return "unchecked", msg # TODO: Call the State Taxation Administration invoice verification API. # This is a placeholder; actual integration requires an enterprise account. return "unchecked", "Tax verification API integration placeholder" ``` The source contains equivalent Chinese comments and messages; the code above is translated into English without changing its behavior. `scripts/compliance_report.py:281-296` ```python def _section_verify_result(records: List[InvoiceRecord]) -> str: """Generate the fourth section: verification results.""" abnormal_records = [ r for r in records if r.verify_status in ("void", "red", "失控", "suspicious", "abnormal") ] if not abnormal_records: return """## IV. Verification Results **Abnormal invoice count**: 0 **Abnormal invoice amount**: ¥0.00 ✅ All invoice statuses are normal. """ ``` The source contains equivalent Chinese report text; it is translated here without changing the vulnerable control flow. ### Technical Analysis The tax verification function does not contact an authoritative tax service. Even for an authorized Pro-tier request, it always returns the `unchecked` state. The report generator only treats explicitly adverse states as abnormal. It does not distinguish between successfully verified normal invoices and invoices for which no verification occurred. If every re ...[truncated 1728 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/compliance_report.py:151
Finding

Untrusted invoice fields are embedded into Markdown and Feishu documents without escaping

Content
View full analysis
str: """Generate the first section: basic information.""" now = datetime.now().strftime("%Y-%m-%d %H:%M:%S") ent_name = buyer_name or "(not provided)" ent_tax_id = buyer_tax_id or "(not provided)" abnormal_count = summary.duplicate_count + summary.suspicious_count rows = [ f"| Item | Content |", f"|------|------|", f"| Report period | {datetime.now().strftime('%Y-%m-%d')} ~ {datetime.now().strftime('%Y-%m-%d')} |", f"| Total invoices | {summary.total_invoices} |", f"| Total amount including tax | {_fmt_currency(summary.total_amount)} |", f"| Abnormal invoices | {abnormal_count} |", f"| Enterprise name | {ent_name} |", f"| Taxpayer identification number | {ent_tax_id} |", ] ``` `scripts/compliance_report.py:254-272` ```python for i, r in enumerate(dup_records, 1): invoice_no = r.invoice_no or "(no number)" date = _fmt_date(r.date) amount = r.amount total_dup_amount += amount seller = r.seller_name or "(no seller)" if r.status == "duplicate": if r.invoice_code and r.invoice_no: reason = "Invoice number is identical" else: reason = "Key fields are identical" else: reason = "Amount, date, and buyer match but invoice number differs" rows.append( f"| {i} | {invoice_no} | {date} | " f"{_fmt_currency(amount)} | {seller} | {reason} |" ) ``` The source contains equivalent Chinese labels and fallback values; these s ...[truncated 2640 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (23)

Vague Triggers

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The trigger list includes broad, common finance-related terms such as '报销', '合规', and '发票识别', which can cause the skill to activate in conversations that only tangentially mention invoices. In this skill’s context, accidental activation is more dangerous because the skill handles sensitive financial/tax data and may steer users toward document parsing, verification, or data export flows without sufficiently specific user intent.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The code is clearly related to invoice processing and does support batch parsing, duplicate checking, XML/OFD handling, and summary generation, so the overall domain matches. However, several core advertised capabilities are not actually implemented in this chunk. The strongest mismatch is authenticity verification: the function explicitly contains a TODO and always returns an unchecked placeholder result. Likewise, the code does not perform OCR or scan recognition; it expects text or raw content to be supplied. The 'report' functionality is limited to aggregate summary JSON, not a compliance report generator tied to the stated accounting circular. Finally, the duplicate detection is deterministic rule logic rather than an identifiable AI-based system. These are material gaps in primary functionality, so this should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
91% confidence
Finding

The declared description presents a broader end-user skill covering invoice ingestion/OCR, duplicate detection, official authenticity verification, compliance reporting, and batch handling. The supplied code chunk is much narrower: its primary function is report generation plus Feishu document/table formatting helpers. It consumes already structured invoice records and summary data, then produces narrative/report output. Duplicate and verification sections depend on precomputed status fields rather than performing those checks. There is no code for file upload, QR/scan/OFD/PDF parsing, OCR, network calls to official verification services, or AI-based analysis. Because major declared core capabilities are absent from the actual code chunk, the description does not accurately represent what this chunk actually does.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The README describes exporting invoice details into Feishu docs and bitable records, but it does not clearly warn users that invoice data contains sensitive corporate and personal financial information that may become shared, retained, or accessible to others in external collaboration systems. Because invoices can include tax IDs, counterparties, dates, and amounts, omission of explicit consent, least-privilege sharing guidance, and retention warnings materially increases data leakage risk.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The trigger phrase list is very broad and can cause the skill to activate for ordinary invoice-related conversations that may not require processing sensitive financial data. In context, this increases the chance that invoice contents, tax IDs, or reimbursement details are unnecessarily routed into parsing, external verification, or reporting workflows.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill describes sending invoice data to the national tax verification platform but does not clearly warn users that invoice identifiers and related details will be transmitted externally for verification. Even if this is expected for the feature, lack of disclosure and consent is risky because users may not realize sensitive billing data is leaving the local workflow.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill encourages exporting invoice details into Feishu documents and tables without a prominent warning or consent checkpoint about sending sensitive financial and tax data to a third-party collaboration platform. Because invoices can contain company names, tax IDs, amounts, and transaction details, unintended disclosure could create confidentiality and compliance risks.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The file is entirely written in Chinese and includes no indication that language selection is optional or constrained by a documented regional requirement. Under the policy, forcing a specific language without user opt-in can be a natural-language locale violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This markdown file is a report template that explicitly includes sensitive business and tax-identifying information such as enterprise name, taxpayer identification number, invoice numbers, dates, and amounts. While it notes the report is for internal compliance reference, it does not warn users about privacy, confidentiality, or careful handling of attached invoice images and structured invoice data.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The document is entirely written in Chinese and describes invoice formats specific to China, but it does not explicitly state that the skill is region-specific or provide any user language/locale choice. Under the policy, forcing a specific language or locale without opt-in or clear justification is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill describes sending invoice fields to an external tax verification endpoint but does not clearly warn users that invoice data will leave the sandbox or obtain explicit consent at the point of transmission. Invoice data can contain sensitive financial and business information, so undisclosed off-box sharing creates a privacy and compliance risk even if the destination is an official tax authority endpoint.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The module header states that Pro/Free permission isolation is implemented, but the CLI initializes TierConfig with is_pro=True by default. This bypasses intended access controls for batch size and tax-verification gating whenever tier data is omitted, allowing unauthorized use of premium behavior and undermining business and control assumptions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This Python file contains the primary module docstring entirely in Chinese, and later user-facing CLI messages are also emitted only in Chinese. Under the policy rule for natural-language violations, forcing a specific language without user opt-in is a locale-policy issue unless the restriction is clearly documented and justified as region-specific.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill claims to perform official invoice authenticity verification, but the implementation always returns an unchecked placeholder and never contacts the tax-platform API. In an invoice-compliance tool, this can cause users or downstream systems to trust invoices as having been verified when no real authenticity check occurred, enabling fraud, duplicate reimbursement, or compliance failures.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This Python file’s natural-language strings, title, docstrings, CLI help, and generated report content are written exclusively in Chinese, indicating the skill is designed to operate in a single language. The file does not offer any user opt-in, language selection, or explicit justification that the locale restriction is required for a region-specific deployment.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The module-level description claims a 'triple-check' including image similarity, but the implementation never performs image-similarity analysis and does not use the image_hash field in duplicate decisions. In a finance/compliance skill, this can create false assurance that visually altered or reformatted duplicate invoices are being detected when they are not.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The code defines tier restrictions and accepts a tier parameter, but check_duplicate ignores the tier entirely, so free/pro access controls are not enforced in this path. If upstream callers rely on this function for entitlement enforcement, users may bypass usage limits or restricted features, undermining billing controls and potentially exposing premium-only processing behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The code embeds Chinese-only field labels and parsing assumptions in natural-language comments and invoice-field definitions, and later relies on Chinese OCR markers for extraction. This imposes a specific locale on skill behavior without offering a language/locale choice or clearly documenting that the skill is intentionally region-specific for compliance or market scope.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

README 标题、说明、流程、FAQ 全部以中文呈现,且没有说明是否仅面向中文用户,或是否支持其他语言/由用户选择输出语言。根据语言/地区政策要求,若技能默认强制单一语言而未提供 opt-in 或明确地域限定,可能构成自然语言策略问题。

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The document content is entirely in Chinese and does not state that the skill is China-specific or that language is configurable. Per the policy, forcing a specific language without user opt-in can be a natural-language policy issue unless the locale constraint is clearly documented and justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
74% confidence
Finding

The file consistently presents instructions and behavior in Chinese only, with no indication that users may choose another language or that the skill is intentionally restricted to a Chinese-speaking or China-specific audience. The policy requires avoiding forced language or locale constraints unless they are opt-in or clearly documented and justified.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The file-level docstring presents this module as an 'InvoiceGuard 合规报告生成器' for generating compliance reports, but a substantial portion of the code also prepares content and schemas specifically for Feishu Docs and Feishu Bitable integrations. While related to reporting, these platform-specific publishing/export behaviors go beyond plain report generation and are not reflected in the module's own stated purpose.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill manifest advertises official tax verification, compliance report generation, batch processing, and invoice upload/scan workflows. In this file, the code only parses invoice text/JSON and returns duplicate-detection results; there is no authenticity-verification API use, report generation, or upload/scan handling here, creating a mismatch if this file is relied on as the described engine.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.