T09 · Insecure Skill Coding Practices
- Location
SKILL.md:110- Finding
Stored HTML Injection in the Public Image Gallery
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 25 and 58, with vulnerable template substitutions at lines 110-118
Vulnerability Type: Stored HTML injection with potential cross-site scripting
Risk Level: HighVulnerable Code
markdown 4. **必须部署画廊**:使用 deploy 工具部署 `/workspace/mj_gallery/`,返回可公开访问的网页链接markdown Step 5: 更新画廊 index.html - 生成新条目加入画廊(最新在前) - 最多保留最近 20 条记录html <div class="card"> <a href="{serial}.webp" target="_blank"> <img src="{serial}.webp" alt="{prompt摘要}"> </a> <div class="info"> <p class="prompt">{prompt中文摘要}</p> <p class="meta">任务: {serial} · MJ 6.1 · {aspect_ratio} · {时间}</p> </div> </div>Technical Analysis
The skill requires an agent to place prompt summaries derived from user input directly into two distinct HTML contexts:
- An attribute context:
alt="{prompt摘要}" - An element-text context:
<p class="prompt">{prompt中文摘要}</p>
The instructions do not require context-sensitive HTML encoding, sanitization, or validation before generating
index.html. An attacker-controlled summary containing quotation marks or HTML metacharacters can therefore terminate the intended attribute or element and inject arbitrary markup. If active HTML attributes or script-capable elements are accepted by the deployment platform, this becomes stored cross-site scripting.The risk is persistent because the generated record is retained in the gallery, and its exposure is broadened by the mandatory deployment to a publicly accessible URL.
Attack Path
- An attacker requests an image using a prompt containing an HTML payload, such as a string intended to close the
altattribute or the prompt paragraph. - The skill translates or summarizes the prompt but preserves enough attacker-controlled markup to form a payload.
- The resulting summary is interpolated into
alt="{prompt摘要}"or `{prompt中文摘 ...[truncated 1064 chars]
- An attribute context:
- Remediation
View remediation
Remediation Suggestions
- Apply context-sensitive HTML encoding to every dynamic value:
- Encode
&,",',<, and>before inserting data into attributes. - Encode
&,<, and>before inserting data into element text.
- Encode
- Prefer a template engine with automatic escaping or safe DOM APIs using
textContentandsetAttributerather than string concatenation. - Treat translated and summarized prompts as untrusted input; translation or summarization is not a security boundary.
- Validate
serialagainst a strict allowlist, such as digits only, before using it in paths or URLs. - Add a restrictive Content Security Policy that disallows inline scripts and event handlers, while recognizing that CSP is defense in depth rather than a substitute for output encoding.
- Add regression tests with payloads targeting both contexts, including:
text "><img src=x onerror=alert(1)> </p><img src=x onerror=alert(1)><p> - Require explicit user approval before public deployment when gallery metadata contains user-supplied content.
- Apply context-sensitive HTML encoding to every dynamic value:
