Back to skill

Security audit

资金与票据技能包(免费版)

Security checks for vulnerabilities and agentic risk

Overview

This skill locally checks user-provided treasury and bill worksheets and I found no hidden network, write, credential, or persistence behavior.

Install only if you are comfortable running a local Node script over exported financial worksheets. Treat results as consistency checks, not audit, legal, approval, or authenticity judgments, and review the reported not-run/withheld checks before relying on the output.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (38)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding
This finding is a weaker variant of the same mismatch theme: the skill appears to act mainly as an orchestrator with structured output rather than literally one line per object, and it discloses withheld checks in the body. While still potentially misleading, this particular discrepancy is less severe than outright claiming unsupported check domains, because the documentation does mention free-version boundaries and non-executed items.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding
This finding is a weaker variant of the same mismatch theme: the skill appears to act mainly as an orchestrator with structured output rather than literally one line per object, and it discloses withheld checks in the body. While still potentially misleading, this particular discrepancy is less severe than outright claiming unsupported check domains, because the documentation does mention free-version boundaries and non-executed items.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding
This finding is a weaker variant of the same mismatch theme: the skill appears to act mainly as an orchestrator with structured output rather than literally one line per object, and it discloses withheld checks in the body. While still potentially misleading, this particular discrepancy is less severe than outright claiming unsupported check domains, because the documentation does mention free-version boundaries and non-executed items.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding
This finding is a weaker variant of the same mismatch theme: the skill appears to act mainly as an orchestrator with structured output rather than literally one line per object, and it discloses withheld checks in the body. While still potentially misleading, this particular discrepancy is less severe than outright claiming unsupported check domains, because the documentation does mention free-version boundaries and non-executed items.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding
This finding is a weaker variant of the same mismatch theme: the skill appears to act mainly as an orchestrator with structured output rather than literally one line per object, and it discloses withheld checks in the body. While still potentially misleading, this particular discrepancy is less severe than outright claiming unsupported check domains, because the documentation does mention free-version boundaries and non-executed items.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding
This finding is a weaker variant of the same mismatch theme: the skill appears to act mainly as an orchestrator with structured output rather than literally one line per object, and it discloses withheld checks in the body. While still potentially misleading, this particular discrepancy is less severe than outright claiming unsupported check domains, because the documentation does mention free-version boundaries and non-executed items.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding
This finding is a weaker variant of the same mismatch theme: the skill appears to act mainly as an orchestrator with structured output rather than literally one line per object, and it discloses withheld checks in the body. While still potentially misleading, this particular discrepancy is less severe than outright claiming unsupported check domains, because the documentation does mention free-version boundaries and non-executed items.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding
This finding is a weaker variant of the same mismatch theme: the skill appears to act mainly as an orchestrator with structured output rather than literally one line per object, and it discloses withheld checks in the body. While still potentially misleading, this particular discrepancy is less severe than outright claiming unsupported check domains, because the documentation does mention free-version boundaries and non-executed items.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding
This finding is a weaker variant of the same mismatch theme: the skill appears to act mainly as an orchestrator with structured output rather than literally one line per object, and it discloses withheld checks in the body. While still potentially misleading, this particular discrepancy is less severe than outright claiming unsupported check domains, because the documentation does mention free-version boundaries and non-executed items.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding
This finding is a weaker variant of the same mismatch theme: the skill appears to act mainly as an orchestrator with structured output rather than literally one line per object, and it discloses withheld checks in the body. While still potentially misleading, this particular discrepancy is less severe than outright claiming unsupported check domains, because the documentation does mention free-version boundaries and non-executed items.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding
This finding is a weaker variant of the same mismatch theme: the skill appears to act mainly as an orchestrator with structured output rather than literally one line per object, and it discloses withheld checks in the body. While still potentially misleading, this particular discrepancy is less severe than outright claiming unsupported check domains, because the documentation does mention free-version boundaries and non-executed items.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding
This finding is a weaker variant of the same mismatch theme: the skill appears to act mainly as an orchestrator with structured output rather than literally one line per object, and it discloses withheld checks in the body. While still potentially misleading, this particular discrepancy is less severe than outright claiming unsupported check domains, because the documentation does mention free-version boundaries and non-executed items.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding
This finding is a weaker variant of the same mismatch theme: the skill appears to act mainly as an orchestrator with structured output rather than literally one line per object, and it discloses withheld checks in the body. While still potentially misleading, this particular discrepancy is less severe than outright claiming unsupported check domains, because the documentation does mention free-version boundaries and non-executed items.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding
This finding is a weaker variant of the same mismatch theme: the skill appears to act mainly as an orchestrator with structured output rather than literally one line per object, and it discloses withheld checks in the body. While still potentially misleading, this particular discrepancy is less severe than outright claiming unsupported check domains, because the documentation does mention free-version boundaries and non-executed items.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding
This finding is a weaker variant of the same mismatch theme: the skill appears to act mainly as an orchestrator with structured output rather than literally one line per object, and it discloses withheld checks in the body. While still potentially misleading, this particular discrepancy is less severe than outright claiming unsupported check domains, because the documentation does mention free-version boundaries and non-executed items.

Ae1

High
Category
analysis-evasion
Content
node scripts/run.mjs --sample # 单对象样例(内置)
Confidence
100% confidence
Finding
Referenced artifact was not completely inspected

Ae1

High
Category
analysis-evasion
Content
node scripts/run.mjs --sample # 单对象样例(内置)
Confidence
100% confidence
Finding
Referenced artifact was not completely inspected

Ae1

High
Category
analysis-evasion
Content
node scripts/run.mjs --sample # 单对象样例(内置)
Confidence
100% confidence
Finding
Referenced artifact was not completely inspected

Ae1

High
Category
analysis-evasion
Content
node scripts/run.mjs --sample # 单对象样例(内置)
Confidence
100% confidence
Finding
Referenced artifact was not completely inspected

Description-Behavior Mismatch

High
Confidence
98% confidence
Finding
The skill metadata promises a broader 14-check treasury/bill verification capability, but this module explicitly implements only a narrow 9-check invoice-consistency subset. In a financial control workflow, that mismatch can cause operators or downstream agents to rely on absent checks and incorrectly treat materials as comprehensively reviewed, creating a dangerous false sense of assurance.

Natural-Language Policy Violations

Medium
Confidence
95% confidence
Finding
The top-level comments and user-facing text throughout the skill are written entirely in Chinese, and returned messages such as advice, findings, notes, and disclaimers are also hard-coded in Chinese. For a general-purpose skill file, this constitutes a natural-language locale constraint without opt-in or a documented region-specific justification.

Description-Behavior Mismatch

Medium
Confidence
98% confidence
Finding
The manifest describes a skill that checks 14 categories across treasury and bill-support documents and outputs one conclusion per object with original file and line references. This module’s own header and code show it handles only '应付账款保理与贴现核对', executes 6 checks, and expects a single pasted text table rather than reviewing the broader document set described in the manifest.

Description-Behavior Mismatch

Medium
Confidence
93% confidence
Finding
The manifest promises that each conclusion includes the original source file and line number. In this file, findings only reference table row numbers derived from the pasted input (for example '第 X 行') and there is no handling of source filenames, original document offsets, or multi-document provenance.

Intent-Code Divergence

Medium
Confidence
90% confidence
Finding
The comment says the parameter is deliberately named `it` to satisfy an external mutation-test pattern for `function checkXxx(it) {`, implying a particular intent and shape. In practice, `checkBlanks` receives and iterates an array of rows rather than a single item, so the inline documentation is actively about gaming a test harness rather than describing the function's real behavior.

Natural-Language Policy Violations

Medium
Confidence
94% confidence
Finding
The file’s natural-language comments, user-facing advice, error text, findings, disclaimer, and sample usage are entirely written in Chinese, indicating the skill is designed to operate in a fixed language/locale. The policy requires flagging language or locale constraints when the skill does not explicitly offer user opt-in or language choice.

Static analysis

No suspicious patterns detected.