T09 · Insecure Skill Coding Practices
- Location
scripts/document_classifier_router.py:327- Finding
Spreadsheet Formula Injection in Exported CSV Files
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This is a local document parsing helper whose file writes are disclosed and aligned with its purpose, but users should handle its generated outputs carefully.
Use this skill only when you intend to turn local documents into stored parsed artifacts. Choose a dedicated output directory, avoid broad batch runs over private folders, treat generated CSV files from untrusted spreadsheets as untrusted, and watch for filename collisions in batch jobs until the publisher adds stronger output isolation.
scripts/document_classifier_router.py:327Spreadsheet Formula Injection in Exported CSV Files
scripts/document_classifier_router.py:391Colliding Document Names Can Overwrite and Mix Parsed Artifacts
The script persistently writes full extracted text to document.md, chunked content to chunks.jsonl, metadata to document.json, and batch summaries to disk, yet this write capability is apparently not declared. In a skill ecosystem, undeclared writes increase the risk of silent data retention, accidental disclosure, and policy bypass because users and orchestrators may not expect sensitive documents to be copied or transformed into searchable artifacts.
The script persistently writes full extracted text to document.md, chunked content to chunks.jsonl, metadata to document.json, and batch summaries to disk, yet this write capability is apparently not declared. In a skill ecosystem, undeclared writes increase the risk of silent data retention, accidental disclosure, and policy bypass because users and orchestrators may not expect sensitive documents to be copied or transformed into searchable artifacts.
The phrase "Use this skill whenever a task involves document parsing, document ingestion, knowledge-base import preparation, or routing files to suitable parsers" casts a very wide net without clear exclusion conditions. This can cause unintended activation for many ordinary document-processing requests, especially since the skill later has a narrower stated scope as a classifier/router.
Listing only "Chinese trigger phrases" establishes a language-specific invocation pattern, but the document does not offer alternative languages or state that Chinese-only activation is intentional for a region-specific use case. This can violate language/locale neutrality expectations for a broadly named 'universal' skill.
The automatic activation section uses expansive examples like document search, RAG corpus construction, and systems that read uploaded documents, which can cause over-selection of this skill in contexts beyond a narrow router role. In agentic systems, ambiguous auto-trigger conditions can lead to inappropriate tool use via exec-enabled parsing flows, increasing the chance of unnecessary file handling or unsafe parser invocation on untrusted inputs.
The integration note instructs automatic invocation for a very broad set of loosely bounded document-related tasks, which can cause the router to be triggered without clear user intent or scope validation. In practice, this increases the chance of processing sensitive files, invoking downstream parsing on untrusted inputs, or performing actions the user did not explicitly authorize.
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
- PDF/Word/PPT/Excel/image parsing
- converting files into standardized parsed outputs
Do not ask the user to remember the full skill name. If you are building or modifying a system that needs document upload, document search, research report retrieval, investment materials management, or knowledge-base enrichment, call this router first.
Canonical CLI path:
The documented commands write parsed outputs to disk and include an option to copy source files, but the integration guidance does not warn users or integrators about these side effects. This can lead to inadvertent data duplication, retention of sensitive source material, or storage of parsed content in locations the user did not expect.
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
Functional local skill package created at:
- `C:\Users\holli\.openclaw\workspace\skills\universal-document-ingestion-router`\n\nThe Skill Workshop proposal exists, but automatic apply failed earlier because the platform reported no approval route. Therefore this folder is the concrete usable skill package.
## Recommended Next Improvements
Dynamic import() can load arbitrary modules at runtime, bypassing static analysis and potentially importing malicious code.
if not have_module(name):
return None
try:
mod = __import__(name)
return getattr(mod, '__version__', None)
except Exception:
return None
The OCR engine is initialized with lang='ch', which forces a specific language/locale behavior for all image and scanned-document processing. This is a natural-language policy issue because the file does not offer a language choice, opt-in, or clear justification for restricting OCR to Chinese.
The manifest describes a 'Document parsing and knowledge-base import router,' which suggests routing decisions for ingestion. In this file, the code goes beyond routing by actually parsing documents, extracting OCR/text/table content, generating markdown documents, chunking content, and writing ingestion artifacts like document.json and chunks.jsonl.
The parser writes full extracted text to markdown and chunked JSONL artifacts, which can persist sensitive document contents beyond the original processing step and make them easier to index, share, or leak. In a knowledge-base import context, this is materially risky because the transformation increases accessibility of confidential content without any visible consent flow or minimization controls.
With --copy-sources enabled, the batch path duplicates original documents into the output directory, potentially consolidating sensitive files into a new location with different permissions, backup policies, or downstream access. This increases data exposure and retention risk, especially in a document-ingestion skill where inputs are likely confidential and users may not realize copying occurs.
The capabilities function enumerates locally installed libraries and executables such as LibreOffice and multiple parsing/OCR packages, including version collection. For a skill described as a document parsing/import router, broad host capability fingerprinting is not clearly justified by the stated purpose, especially when exposed via the CLI 'capabilities' command.
No suspicious patterns detected.