Back to skill

Security audit

Amazon 使用说明改进

Security checks for vulnerabilities and agentic risk

Overview

The skill is not clearly malicious, but it exposes paid ARI account actions, persistent monitoring, exports, and billing-confirmation settings far beyond its narrow instruction-improvement description.

Install only if you trust the ARI service and want this skill to manage more than instruction-improvement reports. Before use, set autoconfirm to ask every time, review any schedule/watch/competitor changes carefully, and treat exports and the local ARI API key as sensitive account data.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (74)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding
The declared purpose is a narrow documentation-improvement skill, but the content instructs the agent to perform many broader actions: paid report generation, account configuration, exports, monitoring, competitor operations, and recurring schedules. This description-behavior mismatch can mislead users and host systems into granting trust or permissions inappropriate for the actual behavior surface.

Ae1

High
Category
analysis-evasion
Content
- CLI:本 Skill 目录下的 `scripts/ari.py`。在 Skill 根目录执行,例如
Confidence
100% confidence
Finding
Referenced artifact was not completely inspected

Ae1

High
Category
analysis-evasion
Content
- CLI:本 Skill 目录下的 `scripts/ari.py`。在 Skill 根目录执行,例如
Confidence
100% confidence
Finding
Referenced artifact was not completely inspected

Ae1

High
Category
analysis-evasion
Content
- CLI:本 Skill 目录下的 `scripts/ari.py`。在 Skill 根目录执行,例如
Confidence
100% confidence
Finding
Referenced artifact was not completely inspected

Ae1

High
Category
analysis-evasion
Content
- CLI:本 Skill 目录下的 `scripts/ari.py`。在 Skill 根目录执行,例如
Confidence
100% confidence
Finding
Referenced artifact was not completely inspected

Ae1

High
Category
analysis-evasion
Content
- CLI:本 Skill 目录下的 `scripts/ari.py`。在 Skill 根目录执行,例如
Confidence
100% confidence
Finding
Referenced artifact was not completely inspected

Ae1

High
Category
analysis-evasion
Content
- CLI:本 Skill 目录下的 `scripts/ari.py`。在 Skill 根目录执行,例如
Confidence
100% confidence
Finding
Referenced artifact was not completely inspected

Ae1

High
Category
analysis-evasion
Content
- CLI:本 Skill 目录下的 `scripts/ari.py`。在 Skill 根目录执行,例如
Confidence
100% confidence
Finding
Referenced artifact was not completely inspected

Ae1

High
Category
analysis-evasion
Content
- CLI:本 Skill 目录下的 `scripts/ari.py`。在 Skill 根目录执行,例如
Confidence
100% confidence
Finding
Referenced artifact was not completely inspected

Ae1

High
Category
analysis-evasion
Content
- CLI:本 Skill 目录下的 `scripts/ari.py`。在 Skill 根目录执行,例如
Confidence
100% confidence
Finding
Referenced artifact was not completely inspected

Ae1

High
Category
analysis-evasion
Content
- CLI:本 Skill 目录下的 `scripts/ari.py`。在 Skill 根目录执行,例如
Confidence
100% confidence
Finding
Referenced artifact was not completely inspected

Ae1

High
Category
analysis-evasion
Content
- CLI:本 Skill 目录下的 `scripts/ari.py`。在 Skill 根目录执行,例如
Confidence
100% confidence
Finding
Referenced artifact was not completely inspected

Ae1

High
Category
analysis-evasion
Content
- CLI:本 Skill 目录下的 `scripts/ari.py`。在 Skill 根目录执行,例如
Confidence
100% confidence
Finding
Referenced artifact was not completely inspected

Context-Inappropriate Capability

High
Confidence
95% confidence
Finding
These commands modify persistent server-side state for monitoring, schedules, competitor relationships, and related review-management workflows, which conflicts with the declared purpose of producing instruction-improvement suggestions only. This is dangerous because an agent or integrator could treat the skill as analytical/read-only while it can actually reconfigure ongoing monitoring and business state on behalf of the user.

Context-Inappropriate Capability

High
Confidence
96% confidence
Finding
Product-operations quote/run/status functionality is unrelated to the declared instruction-improvement purpose and enables execution of broader server-side workflows using account credentials. In a composable agent environment, this creates an unexpected action surface where a skill advertised as analysis-only can trigger operational workflows, increasing the risk of unauthorized business actions or unintended billable processing.

Description-Behavior Mismatch

High
Confidence
98% confidence
Finding
The skill metadata says it is only for extracting instruction-improvement insights from Amazon reviews and not for broader operational actions, but the CLI exposes a much wider administrative surface: account setup, billing-related flows, persistent monitoring, competitor management, exports, workbench state changes, and operations tooling. In an agent setting, this scope mismatch is dangerous because a caller expecting read-only analysis may unknowingly invoke state-changing or cost-incurring commands, violating least privilege and user expectations.

Description-Behavior Mismatch

High
Confidence
97% confidence
Finding
The documented behavior substantially exceeds the manifest’s declared scope, including paid data collection, monitoring, exports, operations workflows, watch features, and report generation. This mismatch can bypass user and platform expectations, increasing the risk of over-privileged use, unintended charges, and hidden access to capabilities that were not disclosed during review or consent.

Natural-Language Policy Violations

Medium
Confidence
93% confidence
Finding
The instructions are written entirely in Chinese and explicitly tell the user what Chinese phrases to send to the AI client, such as the invocation examples on L10 and L17. There is no indication that other languages are supported or that the user can choose their preferred language, which is a natural-language locale policy concern under the stated rule.

Lp3

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding
The skill exposes operational guidance for shell, network, file export, and environment-backed API key usage, but declares no explicit tool/permission scope. That mismatch increases the chance an agent can invoke powerful capabilities beyond the narrow user expectation, including network calls, local file writes, and command execution, without clear containment.

Vague Triggers

Medium
Confidence
93% confidence
Finding
The natural-language trigger conditions are broad enough to match ordinary conversation and cause the agent to infer ASIN/site/range automatically. Over-broad activation can lead to unintended tool use, data access, or billable operations when the user did not intend to invoke this skill specifically.

Autonomous Decision Making

Medium
Category
Excessive Agency
Content
1. 运行 `check`,确认账户、邮箱验证状态和可用积点。
2. 用户要 VOC / 评论分析报告时,默认运行 `voc <ASIN> --site <站点>`。
   **返回里有 `autoConfirmed: true` 就说明已经直接生成了**(1.4.5 起:服务端对前几次小额
   付费操作免确认,用户先拿到结果再谈钱),此时把报告讲给用户,并转述 `autoConfirmNote`
   (本次扣了多少、还剩几次免确认、之后会先问)。**不要在拿到结果后再补问「要不要生成」。**
3. 返回 `confirmationRequired: true` 才需要用户确认:报出 `totalCredits` 与余额,
Confidence
95% confidence
Finding
The skill explicitly tells the agent to run a billable report action by default and accept server-side auto-confirmed charges without first obtaining a fresh user confirmation. Even if the backend permits auto-confirm, this creates a delegated-spend risk where natural-language requests can trigger paid operations the user did not clearly authorize in that session.

Autonomous Decision Making

Medium
Category
Excessive Agency
Content
1. 运行 `check`,确认账户、邮箱验证状态和可用积点。
2. 用户要 VOC / 评论分析报告时,默认运行 `voc <ASIN> --site <站点>`。
   **返回里有 `autoConfirmed: true` 就说明已经直接生成了**(1.4.5 起:服务端对前几次小额
   付费操作免确认,用户先拿到结果再谈钱),此时把报告讲给用户,并转述 `autoConfirmNote`
   (本次扣了多少、还剩几次免确认、之后会先问)。**不要在拿到结果后再补问「要不要生成」。**
3. 返回 `confirmationRequired: true` 才需要用户确认:报出 `totalCredits` 与余额,
   用户同意后运行 `voc <ASIN> --site <站点> --confirm`。该命令会自动补齐采集、等待任务完成、
Confidence
94% confidence
Finding
This workflow continues directly from implicit/default execution into confirmed paid generation logic, making the agent responsible for deciding when to initiate costly operations. That is dangerous because small ambiguities in user intent can result in charges, data collection, and persistent report creation without sufficiently explicit consent boundaries.

Autonomous Decision Making

Medium
Category
Excessive Agency
Content
用户只说自然语言。网页是补充视图(图表、分享链接、海报),不是把人送走的地方。

**确认与扣点**
- 报价返回 `autoConfirm: true` 时直接生成,不要再问「要不要」。生成后一句话交代:本次扣了多少、
  还剩几次免确认(或「免费版小额不问」)。策略由服务端决定:免费版小额不问;付费版前几次不问,之后先问。
- 用户说「以后别问了 / 50 以内直接做」→ 运行 `autoconfirm 50`;说「以后每次先问我」→ `autoconfirm off`;
  说「恢复默认」→ `autoconfirm default`。这是唯一需要你代用户设置的东西,设完复述一句当前规则。
Confidence
96% confidence
Finding
The instruction to proceed automatically whenever `autoConfirm: true` normalizes autonomous spending decisions by the agent. In a skill with paid APIs and report generation, that reduces the user's control over billing and increases the risk of unintended charges from loosely phrased requests.

Autonomous Decision Making

Medium
Category
Excessive Agency
Content
**确认与扣点**
- 报价返回 `autoConfirm: true` 时直接生成,不要再问「要不要」。生成后一句话交代:本次扣了多少、
  还剩几次免确认(或「免费版小额不问」)。策略由服务端决定:免费版小额不问;付费版前几次不问,之后先问。
- 用户说「以后别问了 / 50 以内直接做」→ 运行 `autoconfirm 50`;说「以后每次先问我」→ `autoconfirm off`;
  说「恢复默认」→ `autoconfirm default`。这是唯一需要你代用户设置的东西,设完复述一句当前规则。
- 报价需要确认时,只说两个数:这次多少积点、余额多少,然后等用户一个「好」。采集是**固定单价**:直接说「15 积点/页 × 3 页 = 45 积点」,不要说成「预计 / 最多」——价格不会浮动;商品评论不够这么多页时只收实际采到的页数,差额自动退回(`pricingNote` 已写好这句)。不要罗列参数。
Confidence
90% confidence
Finding
Allowing the agent to change account-level `autoconfirm` settings based on casual phrases like '以后别问了' delegates a persistent billing-control change to natural-language interpretation. Persistent preference changes are more dangerous than one-off actions because they can affect future sessions and future charges.

Autonomous Decision Making

Medium
Category
Excessive Agency
Content
- 报价返回 `autoConfirm: true` 时直接生成,不要再问「要不要」。生成后一句话交代:本次扣了多少、
  还剩几次免确认(或「免费版小额不问」)。策略由服务端决定:免费版小额不问;付费版前几次不问,之后先问。
- 用户说「以后别问了 / 50 以内直接做」→ 运行 `autoconfirm 50`;说「以后每次先问我」→ `autoconfirm off`;
  说「恢复默认」→ `autoconfirm default`。这是唯一需要你代用户设置的东西,设完复述一句当前规则。
- 报价需要确认时,只说两个数:这次多少积点、余额多少,然后等用户一个「好」。采集是**固定单价**:直接说「15 积点/页 × 3 页 = 45 积点」,不要说成「预计 / 最多」——价格不会浮动;商品评论不够这么多页时只收实际采到的页数,差额自动退回(`pricingNote` 已写好这句)。不要罗列参数。

**新手(`check` 返回 `autoConfirm.mode` 为 `first_runs` / `free_small`,或问"然后呢")**
Confidence
88% confidence
Finding
This section reinforces account-level `autoconfirm` manipulation and minimizes detail around billing approval, which can encourage the agent to treat brief acknowledgments as sufficient for consequential financial settings. The danger is not the wording alone but the combination of natural-language triggers, paid actions, and persistent account changes.

Static analysis

No suspicious patterns detected.