Back to skill

Security audit

Universal Document Ingestion Router

Security checks for vulnerabilities and agentic risk

Overview

This is a local document parsing helper whose file writes are disclosed and aligned with its purpose, but users should handle its generated outputs carefully.

Use this skill only when you intend to turn local documents into stored parsed artifacts. Choose a dedicated output directory, avoid broad batch runs over private folders, treat generated CSV files from untrusted spreadsheets as untrusted, and watch for filename collisions in batch jobs until the publisher adds stronger output isolation.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/document_classifier_router.py:327
Finding

Spreadsheet Formula Injection in Exported CSV Files

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/document_classifier_router.py:391
Finding

Colliding Document Names Can Overwrite and Mix Parsed Artifacts

Content
View full analysis
Dict[str, Any]: path = path.resolve() cls = classify(path) out_dir = out_base.resolve() / safe_name(path.stem) out_dir.mkdir(parents=True, exist_ok=True) ``` ```python def batch_process(input_dir: Path, output_dir: Path, copy_sources: bool = False, enable_chunks: bool = True, recursive: bool = True) -> Dict[str, Any]: input_dir = input_dir.resolve() output_dir = output_dir.resolve() parse_root = output_dir / 'parsed' sample_root = output_dir / 'samples' output_dir.mkdir(parents=True, exist_ok=True) results: List[Dict[str, Any]] = [] for source in iter_supported_files(input_dir, recursive=recursive): parse_source = source if copy_sources: sample_root.mkdir(parents=True, exist_ok=True) target = sample_root / source.name if target.exists(): target = sample_root / f"{safe_name(source.stem)}_{sha256_file(source)[:8]}{source.suffix.lower()}" shutil.copy2(source, target) parse_source = target try: parsed = parse_file(parse_source, parse_root, enable_chunks=enable_chunks) ``` The affected artifacts are written later to fixed names within the colliding directory: ```python (out_dir / 'document.md').write_text(md, encoding='utf-8') if enable_chunks: chunks = chunk_text(text, manifest['document_id'], str(path), method) write_jsonl(out_dir / 'chunks.jsonl', chunks) manifest['structure']['chunk_count'] = len(chunks) write_json(out_dir / 'document.json', manifest) ``` ### Technical Analysis The output directory is derived only fro ...[truncated 3109 chars]
Remediation
View remediation
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (15)

Lp1

High
Category
MCP Least Privilege
Confidence
96% confidence
Finding

The script persistently writes full extracted text to document.md, chunked content to chunks.jsonl, metadata to document.json, and batch summaries to disk, yet this write capability is apparently not declared. In a skill ecosystem, undeclared writes increase the risk of silent data retention, accidental disclosure, and policy bypass because users and orchestrators may not expect sensitive documents to be copied or transformed into searchable artifacts.

Content

No source excerpt is available for this finding.

Lp1

High
Category
MCP Least Privilege
Confidence
96% confidence
Finding

The script persistently writes full extracted text to document.md, chunked content to chunks.jsonl, metadata to document.json, and batch summaries to disk, yet this write capability is apparently not declared. In a skill ecosystem, undeclared writes increase the risk of silent data retention, accidental disclosure, and policy bypass because users and orchestrators may not expect sensitive documents to be copied or transformed into searchable artifacts.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The phrase "Use this skill whenever a task involves document parsing, document ingestion, knowledge-base import preparation, or routing files to suitable parsers" casts a very wide net without clear exclusion conditions. This can cause unintended activation for many ordinary document-processing requests, especially since the skill later has a narrower stated scope as a classifier/router.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Listing only "Chinese trigger phrases" establishes a language-specific invocation pattern, but the document does not offer alternative languages or state that Chinese-only activation is intentional for a region-specific use case. This can violate language/locale neutrality expectations for a broadly named 'universal' skill.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The automatic activation section uses expansive examples like document search, RAG corpus construction, and systems that read uploaded documents, which can cause over-selection of this skill in contexts beyond a narrow router role. In agentic systems, ambiguous auto-trigger conditions can lead to inappropriate tool use via exec-enabled parsing flows, increasing the chance of unnecessary file handling or unsafe parser invocation on untrusted inputs.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The integration note instructs automatic invocation for a very broad set of loosely bounded document-related tasks, which can cause the router to be triggered without clear user intent or scope validation. In practice, this increases the chance of processing sensitive files, invoking downstream parsing on untrusted inputs, or performing actions the user did not explicitly authorize.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/agent-integration.md (reported line 16)May include surrounding context.

md
- PDF/Word/PPT/Excel/image parsing
- converting files into standardized parsed outputs

Do not ask the user to remember the full skill name. If you are building or modifying a system that needs document upload, document search, research report retrieval, investment materials management, or knowledge-base enrichment, call this router first.

Canonical CLI path:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The documented commands write parsed outputs to disk and include an option to copy source files, but the integration guidance does not warn users or integrators about these side effects. This can lead to inadvertent data duplication, retention of sensitive source material, or storage of parsed content in locations the user did not expect.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/development-report.md (reported line 122)May include surrounding context.

md
Functional local skill package created at:

- `C:\Users\holli\.openclaw\workspace\skills\universal-document-ingestion-router`\n\nThe Skill Workshop proposal exists, but automatic apply failed earlier because the platform reported no approval route. Therefore this folder is the concrete usable skill package.

## Recommended Next Improvements

Dynamic import via __import__()

Medium
Category
Dangerous Code Execution
Confidence
75% confidence
Finding

Dynamic import() can load arbitrary modules at runtime, bypassing static analysis and potentially importing malicious code.

Content

Scanner excerpt · scripts/document_classifier_router.py (reported line 36)May include surrounding context.

python
if not have_module(name):
        return None
    try:
        mod = __import__(name)
        return getattr(mod, '__version__', None)
    except Exception:
        return None

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The OCR engine is initialized with lang='ch', which forces a specific language/locale behavior for all image and scanned-document processing. This is a natural-language policy issue because the file does not offer a language choice, opt-in, or clear justification for restricting OCR to Chinese.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest describes a 'Document parsing and knowledge-base import router,' which suggests routing decisions for ingestion. In this file, the code goes beyond routing by actually parsing documents, extracting OCR/text/table content, generating markdown documents, chunking content, and writing ingestion artifacts like document.json and chunks.jsonl.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The parser writes full extracted text to markdown and chunked JSONL artifacts, which can persist sensitive document contents beyond the original processing step and make them easier to index, share, or leak. In a knowledge-base import context, this is materially risky because the transformation increases accessibility of confidential content without any visible consent flow or minimization controls.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

With --copy-sources enabled, the batch path duplicates original documents into the output directory, potentially consolidating sensitive files into a new location with different permissions, backup policies, or downstream access. This increases data exposure and retention risk, especially in a document-ingestion skill where inputs are likely confidential and users may not realize copying occurs.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The capabilities function enumerates locally installed libraries and executables such as LibreOffice and multiple parsing/OCR packages, including version collection. For a skill described as a document parsing/import router, broad host capability fingerprinting is not clearly justified by the stated purpose, especially when exposed via the CLI 'capabilities' command.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.