Back to skill

Security audit

财务月度自查包(免费版)

Security checks for vulnerabilities and agentic risk

Overview

The local checker code is mostly offline, but the skill also pushes an external paid purchase and installation flow despite claiming no payment and no network.

Install only if you want a Chinese-language local checker for the three disclosed free sections. Treat the paid-upgrade instructions separately: do not let an agent follow external purchase, raw IP, QR-code, or installation links unless you intentionally want to buy the paid product, have verified the merchant and product through trusted platform channels, and your platform allows that flow. Do not rely on it as a general financial reconciliation or audit tool.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (19)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding
整体方向是财务分段核对,本地执行、材料不足不给结论,这些与描述基本一致。但存在实质性范围差异:代码不是通用核对引擎,而是一个受限的月度自查包编排器,只运行3个硬编码检查项;同时描述中的“重复检测”和“每条结论引用原文”在当前代码块中没有对应实现。代码还显式列出部分检查项 withheld/not run,与描述中较宽泛的“免费版执行引擎声明的免费检查项”相比,实际能力更窄。因此应判定为描述与行为存在明显不匹配。

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding
代码没有发现越权行为、外部网络访问或隐藏资源使用,权限层面基本一致;问题在于“描述范围”和“实际行为范围”不一致。描述把技能表述成通用的分段财务核对材料检查,而代码实际专门检查应付账款账龄与付款计划台账,并且只支持识别特定列名和固定6个免费检查项。此外,描述承诺“每条结论引用原文”,但代码仅返回按行号定位的结构化消息,没有抽取并附带原文引用。因此应判定为描述与行为存在实质性不匹配。

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding
该代码的主用途和适用对象比声明显著更窄且更具体。代码注释、service_type(CIT_PREPAY_CHECK)、固定列映射、计算公式和校验逻辑都表明其仅用于企业所得税季度预缴表的算术/勾稽检查,不是通用的“分段财务核对材料逐项核对”引擎。虽然声明提到的若干检查类型(逐行复算、合计、重复、空缺)与代码部分重合,但缺失了关键限定:代码必须是特定税表文本,且并未提供“每条结论引用原文”的能力。另有场景不一致:声明触发词偏月度/泛财务自查,代码却聚焦季度企业所得税预缴。因此应判定为描述与实际行为存在实质性不匹配。

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding
该代码的主目的非常明确:解析差旅报销表,并做差旅费标准/报销一致性检查。它识别的列名、检查项、输出 service_type、示例文本、说明文档都围绕差旅费(住宿、交通、伙食补助、审批等)展开。虽然声明中提到的部分免费检查项如逐行复算、合计勾稽、重复与空缺检测,在形式上与代码中的部分检查相似,但声明将技能描述为更广义的“分段财务核对材料”工具,这会让人预期其能处理更一般的财务核对材料,而代码只支持差旅报销这一狭窄场景。此外,声明特别强调“每条结论引用原文”,代码并未提取原文片段引用,只给出行号、字段名和复算结果。因此这是实质性的描述与行为不匹配。

Intent-Code Divergence

High
Confidence
98% confidence
Finding
The file makes strong no-network and fully-local assurances, then later instructs use of external websites and payment endpoints. Such contradictory claims are dangerous because they reduce user caution and may cause operators or automated systems to grant trust or permissions they would not otherwise allow.

Ae1

High
Category
analysis-evasion
Content
node scripts/run.mjs --sample # 先看样例
Confidence
100% confidence
Finding
Referenced artifact was not completely inspected

Ae1

High
Category
analysis-evasion
Content
node scripts/run.mjs --sample # 先看样例
Confidence
100% confidence
Finding
Referenced artifact was not completely inspected

Description-Behavior Mismatch

High
Confidence
97% confidence
Finding
The skill claims it is fully local and requires no network, yet embeds instructions that push the agent into external purchase and installation flows through multiple URLs. This hidden expansion of behavior breaks user trust and can cause the agent to leave the local-only boundary, exposing users to phishing, unsafe installs, or unwanted commercial actions.

Context-Inappropriate Capability

High
Confidence
96% confidence
Finding
Embedding payment and product-installation workflows inside a financial-checking skill creates an unjustified capability crossover. A user invoking a bookkeeping check should not implicitly expose themselves to transaction initiation or software installation behavior, which materially raises the risk of social engineering and supply-chain abuse.

Description-Behavior Mismatch

High
Confidence
99% confidence
Finding
The implementation is materially misaligned with the declared skill purpose: instead of a general monthly segmented financial self-check, it performs a narrow travel-expense reimbursement check. This can cause users or orchestrators to send the wrong data to the skill and rely on incomplete or irrelevant validation results, creating a security/integrity risk through false assurance and scope confusion.

Natural-Language Policy Violations

Medium
Confidence
90% confidence
Finding
The file’s operative instructions to the agent are written entirely in Chinese and mandate exact user-facing messaging and purchase-flow disclosures, but they do not offer the user a language or locale choice. This can violate a language/locale policy when the skill is used in multilingual contexts, especially because the instructions are framed as mandatory behavior for every invocation.

Whitespace Padding

Medium
Category
Prompt Injection
Content
输出与你的安全约束**;它**可被用户关闭**(`SKILL_NO_UPSELL=1`),也**必须可忽略**(不点不买不影响结果)。
   执行官方购买指令时:先读一次官方说明;核对**商户号 `2088702818190040`**、商品号与订单是否一致;
   ⛔ 不要点击来路不明的链接/二维码,也**不要替用户付款** —— 付款一律由**用户本人**确认。
Confidence
80% confidence
Finding
Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Ssd 4

Medium
Confidence
93% confidence
Finding
The step-by-step narrative is designed to have the agent obtain authorization, initiate purchase steps, and continue installation on the user's behalf. Even though it says the user must confirm payment, this still normalizes transactional delegation and can be abused to steer the agent into risky external actions and installation of unreviewed dependencies.

Natural-Language Policy Violations

Medium
Confidence
95% confidence
Finding
The file’s natural-language instructions, sample input expectations, labels, and user-facing messages are entirely in Chinese, including required table headers and remediation text. There is no indication that the skill is region-specific or that users may opt into another language, which makes this a language/locale policy concern.

Natural-Language Policy Violations

Medium
Confidence
96% confidence
Finding
This JavaScript file contains extensive Chinese-only comments and user-facing output such as advice, error messages, notes, and disclaimers. Under the natural-language policy rule, forcing a specific language without opt-in is a locale policy violation unless the constraint is explicitly documented and justified as region-specific.

Natural-Language Policy Violations

Medium
Confidence
96% confidence
Finding
This JavaScript file contains natural-language instructions, sample data, field labels, and user-facing error/disclaimer text entirely in Chinese, including required input guidance and output messages. Because the skill does not offer any language selection or document a justified locale restriction, it appears to enforce a specific language contrary to the language/locale policy.

Intent-Code Divergence

Medium
Confidence
95% confidence
Finding
The disclaimer states that a meal-allowance formula check is performed, but the free-tier code explicitly withholds that logic and never executes it. This is dangerous because it can mislead users into trusting a control that does not exist, allowing over-limit or inconsistent meal claims to pass without review.

Intent-Code Divergence

Low
Confidence
91% confidence
Finding
The manifest and surrounding documentation consistently describe this skill as handling segmented financial reconciliation materials. The comment on L077, however, says that invalid JSON will be treated as plain-text material and gives "contract full text" as the example, which points to a different document domain and contradicts the intended skill scope.

Natural-Language Policy Violations

Low
Confidence
98% confidence
Finding
The JSON value is entirely Chinese-language business content, indicating a fixed locale/language presentation. For all file types, a skill should not force a specific language unless it offers user opt-in or clearly documents a justified locale restriction.

Static analysis

No suspicious patterns detected.