T09 · Insecure Skill Coding Practices
- Location
scripts/generate_report.py:958- Finding
Stored HTML Injection Through Unescaped Model Names
- Content
View full analysis
Vulnerability Details
File Location:
scripts/generate_report.py, lines 958–961 and 1021–1023
Vulnerability Type: Stored HTML injection / cross-site scripting
Risk Level: MediumVulnerable Code
python json_lines = [] for m in unconfigured: json_lines.append(f' "{m}": {{') json_lines.append(f' "input": 0, // Fill in the actual input price') json_lines.append(f' "output": 0 // Fill in the actual output price') json_lines.append(f" }},")The resulting strings are later inserted into the HTML report without escaping:
python out.append(' <pre><code class="language-json">') out.extend(json_lines) out.append("</code></pre>")Technical Analysis
The report generator builds a sample JSON configuration using names from
meta.unconfigured_models. These model identifiers originate from automatically ingested session and trace data.Although the model name is escaped with
_esc(m)in the adjacent HTML table, it is interpolated directly intojson_lines. Those lines are then placed inside an HTML<pre><code>element without HTML encoding. The code element does not neutralize markup; an identifier containing closing tags can terminate the surrounding elements and inject active HTML or JavaScript.A malicious model identifier such as:
html </code></pre><script>/* attacker-controlled JavaScript */</script><pre><code>would therefore become executable markup when a generated HTML report is opened.
Attack Path
- An attacker-controlled model provider, imported session record, or other party able to influence automatically ingested trace data supplies a crafted model identifier.
- The collector records that identifier as an unconfigured model.
- The user generates an HTML report from the collected data.
_build_unconfigured_models_section()interpolates the identifier intojson_lines.- The HTML renderer inserts those lines directly into a
<pre><code>bloc ...[truncated 891 chars]
- Remediation
View remediation
Remediation Suggestions
- Serialize the configuration example with
json.dumps()rather than manually constructing JSON-like strings. - Apply
_esc()orhtml.escape(..., quote=True)to the complete serialized snippet before inserting it into HTML:
python snippet = json.dumps(stub, ensure_ascii=False, indent=4) out.append(' <pre><code class="language-json">') out.append(_esc(snippet)) out.append(" </code></pre>")- Keep Markdown and HTML rendering separate so escaping appropriate to each output context is always applied.
- Add an HTML-injection regression test using an unconfigured model name containing:
html </code></pre><script>alert(1)</script>The test should assert that the generated report contains only encoded forms such as
<script>and no executable<script>element. 5. Review all other HTML construction paths for values originating from traces, session metadata, pricing endpoints, imported files, and model names. Require contextual encoding for text nodes, attributes, URLs, and script contexts. 6. Consider a restrictive Content Security Policy for generated HTML as defense in depth, while retaining correct output encoding as the primary fix.- Serialize the configuration example with
