T09 · Insecure Skill Coding Practices
- Location
SKILL.md:33- Finding
Unbounded Processing of Untrusted PPTX Archives
- Content
View full analysis
]*>//g' ``` ```bash # Lines 33-44 unzip -p "file.pptx" "ppt/slides/slide*.xml" 2>/dev/null | sed 's/<[^>]*>//g' | tr -s ' \n' cd /path/to/ppt/ for i in {1..N}; do echo "=== Slide $i ===" unzip -p "file.pptx" "ppt/slides/slide$i.xml" 2>/dev/null | sed 's/<[^>]*>//g' | tr -s ' \n' echo "" done ``` ```bash # Lines 102-110 PAGE_COUNT=$(unzip -l "$PPT_FILE" | grep -c "slide[0-9]*\.xml") echo "Total slides: $PAGE_COUNT" for i in $(seq 1 $PAGE_COUNT); do echo "=== Slide $i ===" unzip -p "$PPT_FILE" "ppt/slides/slide$i.xml" 2>/dev/null | sed 's/<[^>]*>//g' | tr -s ' \n' echo "" done ``` ### Technical Analysis The documented workflow processes user-supplied PPTX files directly as ZIP archives without first enforcing limits on: - Compressed file size - Total expanded size - Compression ratio - Number of archive entries - Number of reported slides - Maximum size of an individual slide XML file - Execution time or generated output volume The `unzip -p` command streams the complete expanded contents of matching archive entries into subsequent commands. Piping the data through `sed` and `tr` does not restrict decompression or output size. A highly compressed slide XML entry can therefore expand into a very large stream. The page-count workflow also trusts the number of archive entries matching the slide filename pattern and passes that count to `seq`. An archive with an excessive number of matching entries can cause a large loop and repeated `unzip` process creation, including attempts to read slide names that may not exist sequentially. Because extraction uses `unzip -p` rather than writing archive paths to ...[truncated 1676 chars]- Remediation
View remediation
/dev/null`. - Treat malformed archives, duplicate entries, metadata inconsistencies, and exceeded limits as explicit validation failures. - Return a concise error rather than continuing partial or repeated extraction. 6. **Use a hardened parser where possible** - Prefer a PPTX-processing library that supports controlled ZIP entry access and explicit size limits. - Parse slide XML with an XML parser configured to disable external entities and external resource resolution. ]]>
