T09 · Insecure Skill Coding Practices
- Location
scripts/docx_extractor.py:24- Finding
Unbounded Processing of Attacker-Controlled Office Archives
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill is an offline Office-to-Markdown converter with a resource-use risk on large or malicious documents, but no evidence of hidden access, network activity, persistence, or exfiltration.
Reasonable to install for local document conversion. Use caution with unknown or very large Office files, and prefer running conversions in a constrained environment if processing untrusted uploads because the parser does not enforce size, decompression, or output limits.
scripts/docx_extractor.py:24Unbounded Processing of Attacker-Controlled Office Archives
The code is consistent with part of the description: it uses only the Python standard library, works offline, performs no network access, and converts DOCX content to Markdown. However, the declared purpose materially overstates the supported formats by claiming DOCX, XLSX, and PPTX support. This code chunk only processes DOCX files and contains no spreadsheet or PowerPoint handling. That makes the description inaccurate relative to the supplied code chunk.
Yes, this is a mismatch. The declared description promises a full offline Office document to Markdown converter handling Word, Excel, and PowerPoint files. The supplied code does not implement any such functionality; it only exposes a local xmlfile symbol and package metadata for et_xmlfile. There is no evidence in this chunk of reading Office files, extracting content, converting anything to Markdown, or orchestrating document conversion. The code’s apparent purpose is package initialization for an XML-related helper library, which is materially different from the declared primary purpose.
The supplied code is a low-level XML serialization module derived from ElementTree. Its functions write ElementTree structures to XML/HTML/text with custom namespace handling (IncrementalTree.write, _serialize_ns_xml, tostring, tostringlist, etc.). While such code could be a supporting dependency inside a larger Office-processing package, this chunk by itself does not implement the declared skill purpose of converting Word/Excel/PowerPoint documents to Markdown or extracting content from those formats. It only handles generic XML tree serialization and file writing, operating offline with standard library modules. Therefore the description materially overstates and misrepresents this code chunk's actual behavior.
The supplied code chunk is an internal XML writer implementation from et_xmlfile/openpyxl-style infrastructure. It manages XML element context, namespace handling, and incremental serialization to a file-like output. This is materially different from the declared purpose of converting Office documents to Markdown. While XML handling could be a supporting building block in a larger Office-processing package, this specific code neither reads Office files nor extracts text nor emits Markdown. Therefore, the description does not accurately represent what this code chunk actually does.
The declared description promises a broad Office-to-Markdown converter covering Word, Excel, and PowerPoint, implemented purely in Python and usable offline. The supplied code chunk, however, is only openpyxl's init.py, which imports and exposes workbook-related Excel loading functionality and package constants. There is no evidence here of Markdown rendering, text extraction pipelines, or support for DOCX and PPTX. While openpyxl is consistent with part of the Excel-processing claim and does not show suspicious external access, the primary purpose represented by this code chunk is materially narrower and different from the declared skill description.
This code chunk does not match the declared purpose. The description claims a functional offline converter for Office documents to Markdown, but the actual code only contains static metadata for openpyxl. Metadata files are not supporting implementation for conversion behavior by themselves; they provide no evidence of parsing Word, Excel, or PowerPoint files or producing Markdown. Therefore, the code’s actual behavior is materially different from the declared purpose.
The declared description promises an offline pure-Python converter from Office formats (DOCX/XLSX/PPTX) to Markdown. The supplied code does something materially different: it prepares and writes XML elements for spreadsheet cells as part of openpyxl's XLSX-writing machinery. It sets cell attributes, converts date values to Excel/ISO formats, handles formulas and rich text, and appends hyperlinks to worksheet metadata. There is no Markdown generation, no document text extraction, no DOCX or PPTX support, and no general file conversion logic. Although it is offline and pure Python, those are incidental properties and do not make the implementation match the declared primary purpose.
The supplied code chunk is a library module for representing and managing individual spreadsheet cells in Excel workbooks. It validates strings, infers Excel data types, manages date formats, hyperlinks, comments, and merged-cell behavior. This does not implement conversion of Office files to Markdown, nor does it process Word or PowerPoint documents. While it is related to XLSX internals, its actual purpose is materially different from the declared end-user capability, so the description does not accurately represent the code.
The declared description promises a complete offline converter for DOCX/XLSX/PPTX to Markdown. The supplied code chunk instead implements internal read-only cell objects used by openpyxl for spreadsheet handling. It exposes cell coordinates, values, styles, and date/format helpers, plus an EmptyCell sentinel. This is a narrow, low-level support component for XLSX reading, not a converter and not a cross-format Office extraction tool. While it is consistent with part of XLSX processing and is pure Python/offline, the primary purpose is materially different from the declared skill behavior.
The declared description promises a complete offline converter for DOCX/XLSX/PPTX to Markdown. However, this code chunk only implements internal rich text data structures for spreadsheet cells, specifically for openpyxl. It validates elements, constructs rich text objects, parses from XML nodes, optimizes adjacent text runs, and serializes back to XML. There is no logic for reading Office files broadly, no Markdown generation, and no support for Word or PowerPoint documents. This is a material mismatch in primary purpose and capability, not just an implementation detail.
The declared description promises a complete offline converter for DOCX, XLSX, and PPTX into Markdown. The supplied code chunk instead contains internal data-model classes from openpyxl for representing rich text and phonetic properties in Excel cell text. Its only notable behavior is storing text-formatting metadata and exposing a content property that joins text fragments. There is no code for opening Office documents generally, traversing document structure, extracting content from Word/PowerPoint, or emitting Markdown. This is therefore a clear description-behavior mismatch in primary purpose and capabilities.
The declared description says the skill converts DOCX/XLSX/PPTX documents to Markdown offline. The supplied code chunk instead is a small part of openpyxl's chart model for representing 3D chart settings in Excel files. It only defines data structures and validation for chart view/surface properties; it does not read Office documents for extraction, convert any format to Markdown, or process Word/PowerPoint files at all. This is a clear purpose mismatch, not merely an implementation detail.
The declared description says the skill converts Office documents to Markdown offline. The provided code only imports and exposes chart-related classes from openpyxl, specifically for Excel chart types and references. There is no conversion pipeline, no parsing of Word/PowerPoint documents, no spreadsheet-to-Markdown rendering, and no text extraction behavior. This is a materially different primary purpose, so it is a clear mismatch.
The declared description says the skill converts Office documents (Word, Excel, PowerPoint) to Markdown offline. However, this code is specifically an internal chart module from openpyxl focused on representing and writing Excel chart objects. Its methods add chart data, set categories, manage axes and legends, and serialize chart data to XML. There is nothing here that extracts document text, processes DOCX/PPTX, or produces Markdown output. This is a clear description-to-behavior mismatch with a materially different primary purpose.
The declared description says the skill converts Microsoft Office files (DOCX, XLSX, PPTX) to Markdown offline in pure Python. The supplied code chunk does something entirely different: it is a library module from openpyxl that defines AreaChart and AreaChart3D classes and related chart metadata/axes behavior for Excel workbook chart structures. There is no file I/O, no document parsing, no extraction of text content, and no Markdown generation. This is a material mismatch in primary purpose and actual capability.
The declared description says the skill converts Office documents (Word, Excel, PowerPoint) to Markdown without dependencies. However, this code is a narrow internal library module from openpyxl dealing with chart axis representations and XML serialization for Excel charts. It defines classes like NumericAxis, TextAxis, DateAxis, and SeriesAxis, along with related formatting/scaling metadata. There is nothing here for reading DOCX/PPTX, extracting document text, or producing Markdown output. This is a materially different primary purpose, so the description does not accurately represent the supplied code chunk.
This code chunk is part of openpyxl's chart subsystem, specifically classes representing bar charts and 3D bar charts in Excel workbooks. It defines schema-like chart attributes such as grouping, axes, gap width, overlap, and legends. There is no logic for reading Office files broadly, no Markdown generation, and no document text extraction. The declared description presents a document conversion utility covering DOCX, XLSX, and PPTX, but the actual code is narrowly focused on Excel chart object definitions, making the primary purpose materially different.
The supplied code chunk is an autogenerated openpyxl module defining a BubbleChart class and its descriptors. It models Excel bubble chart metadata and axes; it does not parse Office files, extract text, or emit Markdown. There is no conversion logic for DOCX, XLSX, or PPTX in this snippet. Therefore the code’s actual behavior is materially different from the declared skill description.
The declared description says the skill converts Office documents to Markdown without dependencies. This code chunk does not implement document conversion, text extraction, or Markdown output. It is a library module from openpyxl that defines serializable data structures for Excel chart spaces and related properties, including XML namespace handling. Its primary purpose is chart representation/serialization inside spreadsheet processing, which is materially different from the declared skill purpose.
The declared purpose describes a broad offline converter for Word, Excel, and PowerPoint documents into Markdown. The actual code is a narrow internal library module from openpyxl for representing chart data sources in Excel files. It contains class definitions and validation descriptors for numeric/string chart references and cached values, with no parsing of Office documents into Markdown, no extraction pipeline, and no support for DOCX or PPTX. This is a material description-behavior mismatch because the code’s primary purpose is unrelated to the claimed conversion functionality.
The declared purpose describes a complete offline Office-to-Markdown extraction/conversion capability across Word, Excel, and PowerPoint files. The actual code chunk is a small internal utility module from openpyxl related to chart descriptors and number formatting. It does not read Office files, extract text, process DOCX/PPTX, or emit Markdown. This is a materially different primary purpose, so the description does not accurately represent the supplied code.
The supplied code is a narrow internal library component from openpyxl for representing Excel chart error bar settings. It declares fields like error direction, bar type, value type, plus/minus data sources, and graphical properties. There is no logic for reading Office files, extracting content, converting anything to Markdown, or handling DOCX/PPTX at all. The primary purpose is materially different from the declared description, so this is a clear mismatch.
The supplied code is a small internal module from openpyxl that models chart data labels for Excel serialization/deserialization. It contains class definitions and field descriptors only; there is no logic to open Office files, extract document contents, traverse Word/PowerPoint structures, read spreadsheet cells for export, or emit Markdown. This is materially different from the declared purpose of converting Office documents to Markdown. The mismatch is strong because the code’s primary purpose is unrelated object modeling for chart labels, not document conversion.
The supplied code chunk is a small internal component from openpyxl that models chart layout settings for spreadsheets. It declares serializable classes and field constraints for chart positioning and sizing. There is no code to open Office files, extract document contents, convert anything to Markdown, or process DOCX/PPTX/XLSX documents broadly. This is a materially different purpose from the declared skill description, so it is a clear mismatch.
The declared description claims a broad Office-document-to-Markdown conversion capability covering DOCX, XLSX, and PPTX extraction offline. The supplied code chunk does not implement conversion, text extraction, Markdown generation, file parsing, or any end-user workflow. Instead, it is a small internal model from openpyxl representing Excel chart legend elements for serialization. This is a materially different primary purpose, so the description does not accurately represent the code.
No suspicious patterns detected.