T09 · Insecure Skill Coding Practices
- Location
scripts/docx_to_html.py:29- Finding
Stored HTML Injection in DOCX-to-HTML Conversion
- Content
View full analysis
{text}" elif "heading 2" in style: return f"{text}
" elif "heading 3" in style: return f"{text}
" elif "list" in style: return f"- {text}
" else: return f"{text}
" ``` ### Technical Analysis Paragraph text extracted from a potentially untrusted DOCX file is interpolated directly into HTML elements without HTML encoding or sanitization. Consequently, characters such as `<`, `>`, `&`, single quotes, and double quotes retain their markup semantics. An attacker can place active HTML in a DOCX paragraph, for example: ```html``` The converter writes this content unchanged inside a generated paragraph: ```html
``` When the generated file is opened in a browser or published by a web service, the event handler executes as JavaScript. Other payloads may create deceptive forms, initiate browser requests, alter displayed content, or interact with data available to the generated document's origin. The same unsafe conversion pattern is also documented in the inline example at `SKILL.md:232-251`, which may cause users to reproduce the vulnerability even if the standalone script is corrected. ### Attack Path 1. An attacker creates a DOCX document containing HTML or JavaScript-bearing markup in a paragraph. 2. The victim or an automated service runs: ```bash python scripts/docx_to_html.py attacker.docx output.html ``` 3. `para_to_html()` reads the attacker-controlled paragraph text and inserts it directly into an ...[truncated 1170 chars]- Remediation
View remediation
{text}" elif "heading 2" in style: return f"{text}
" elif "heading 3" in style: return f"{text}
" elif "list" in style: return f"- {text}
" return f"{text}
" ``` Additional hardening measures: 1. Apply contextual encoding to every untrusted value written into HTML, including any future attributes, links, image names, table cells, comments, or metadata. 2. If preserving limited document formatting requires accepting HTML, process it through a mature allowlist-based sanitizer rather than interpolating it directly. 3. Disallow scripts, inline event handlers, unsafe URL schemes, active SVG content, embedded objects, and dangerous CSS. 4. Serve generated HTML from an isolated origin with no sensitive cookies or application data. 5. Apply a restrictive Content Security Policy where generated content is hosted, such as disabling scripts unless they are explicitly required. 6. Add regression tests covering script tags, event-handler attributes, malformed tags, SVG payloads, entity-encoded payloads, and quotation-mark breakout attempts. 7. Correct the corresponding unsafe inline implementation in `SKILL.md:232-251`. ]]>
