Back to skill

Security audit

Agent 用量分析报告

Security checks for vulnerabilities and agentic risk

Overview

This analytics skill is mostly coherent, but it needs Review because it reads sensitive local agent histories and has under-disclosed local config writes plus an HTML report injection flaw.

Install only if you are comfortable letting the skill read local agent usage histories and generate reports that may contain private prompts, project paths, costs, task labels, and file names. Treat HTML reports as sensitive, avoid opening or sharing reports built from untrusted/imported logs until the unconfigured-model escaping bug is fixed, and only allow pricing.local.json or pricing.json changes after reviewing the exact models and prices being written.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/generate_report.py:958
Finding

Stored HTML Injection Through Unescaped Model Names

Content
View full analysis

Vulnerability Details

File Location: scripts/generate_report.py, lines 958–961 and 1021–1023
Vulnerability Type: Stored HTML injection / cross-site scripting
Risk Level: Medium

Vulnerable Code

python
json_lines = []
for m in unconfigured:
    json_lines.append(f'    "{m}": {{')
    json_lines.append(f'        "input": 0,   // Fill in the actual input price')
    json_lines.append(f'        "output": 0   // Fill in the actual output price')
    json_lines.append(f"    }},")

The resulting strings are later inserted into the HTML report without escaping:

python
out.append('            <pre><code class="language-json">')
out.extend(json_lines)
out.append("</code></pre>")

Technical Analysis

The report generator builds a sample JSON configuration using names from meta.unconfigured_models. These model identifiers originate from automatically ingested session and trace data.

Although the model name is escaped with _esc(m) in the adjacent HTML table, it is interpolated directly into json_lines. Those lines are then placed inside an HTML <pre><code> element without HTML encoding. The code element does not neutralize markup; an identifier containing closing tags can terminate the surrounding elements and inject active HTML or JavaScript.

A malicious model identifier such as:

html
</code></pre><script>/* attacker-controlled JavaScript */</script><pre><code>

would therefore become executable markup when a generated HTML report is opened.

Attack Path

  1. An attacker-controlled model provider, imported session record, or other party able to influence automatically ingested trace data supplies a crafted model identifier.
  2. The collector records that identifier as an unconfigured model.
  3. The user generates an HTML report from the collected data.
  4. _build_unconfigured_models_section() interpolates the identifier into json_lines.
  5. The HTML renderer inserts those lines directly into a <pre><code> bloc ...[truncated 891 chars]
Remediation
View remediation

Remediation Suggestions

  1. Serialize the configuration example with json.dumps() rather than manually constructing JSON-like strings.
  2. Apply _esc() or html.escape(..., quote=True) to the complete serialized snippet before inserting it into HTML:
python
snippet = json.dumps(stub, ensure_ascii=False, indent=4)
out.append('            <pre><code class="language-json">')
out.append(_esc(snippet))
out.append("            </code></pre>")
  1. Keep Markdown and HTML rendering separate so escaping appropriate to each output context is always applied.
  2. Add an HTML-injection regression test using an unconfigured model name containing:
html
</code></pre><script>alert(1)</script>

The test should assert that the generated report contains only encoded forms such as &lt;script&gt; and no executable <script> element. 5. Review all other HTML construction paths for values originating from traces, session metadata, pricing endpoints, imported files, and model names. Require contextual encoding for text nodes, attributes, URLs, and script contexts. 6. Consider a restrictive Content Security Policy for generated HTML as defense in depth, while retaining correct output encoding as the primary fix.

Vulnerability Patterns
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (180)

Tp4

High
Category
MCP Tool Poisoning
Confidence
91% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
91% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

A skill positioned as read-only analytics also includes workflows for validating, updating, backing up, and applying pricing configuration changes. Bundling configuration maintenance with reporting can mislead users into granting broader capabilities than necessary and creates opportunities for unwanted local state changes.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

A top-level read-only/privacy claim that conflicts with later instructions to write local pricing files is especially dangerous because users and hosting platforms may rely on the summary claim when deciding whether to trust or permit the skill. This kind of capability understatement can lead to unauthorized local modifications under a false safety model.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 288)May include surrounding context.

md
- 每条 pattern 有权重(`scripts/task_rules.json` 可编辑):强信号(如 `skillhub install`、`SKILL.md`、`imagegen`)权重 2~2.5,弱信号(如「了解」「对比」)权重 0.5~0.8,需多条叠加才能取胜。

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

This is a real integrity/scope violation: the skill is described as read-only analytics, but the documentation instructs the agent to enumerate local data and then modify a local JSON file. Even though the write target is framed as a user-local override and requires user participation for pricing values, it expands the skill from passive reporting into filesystem mutation, which can surprise users and normalize post-download agent-driven file edits.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
tests/conftest.py:25

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
tests/test_channel_attribution.py:92