T09 · Insecure Skill Coding Practices
- Location
main.py:319- Finding
Unbounded In-Memory Processing of Untrusted Office Files
- Content
View full analysis
Dict[str, Any]: """处理数据清洗""" data = params.get("data") if isinstance(data, str): data = base64.b64decode(data) request = DataCleaningRequest( data=data, remove_duplicates=params.get("remove_duplicates", True), handle_missing=params.get("handle_missing", "mean"), remove_outliers=params.get("remove_outliers", True), outlier_method=params.get("outlier_method", "iqr"), ) result = await self.sheet_processor.clean_data(request) ``` The same unrestricted decoding pattern is present in the spreadsheet analysis, chart-generation, and PDF handlers. PDF merge also decodes every supplied file into a list before parsing: ```python pdf_files = [ base64.b64decode(file_data) for file_data in params.get("files", []) ] ``` The downstream processors parse XLSX and PDF content synchronously and in memory. XLSX files are ZIP-based containers, while PDFs can contain highly compressed or parser-expensive structures. Consequently, a comparatively small request can cause disproportionate CPU or memory consumption. The configured maximum execution time does not itself establish memory, decoded-size, page-count, or archive-expansion limits. ### Attack Path 1. An attacker invokes a spreadsheet or PDF operation exposed by `execute()`. 2. The attacker supplies an extremely large Base64 value, numerous PDF merge inputs, or a compact but parser-expensive document. 3. The Skill decodes the co ...[truncated 665 chars]- Remediation
View remediation
