T09 · Insecure Skill Coding Practices
- Location
scripts/image_to_markdown.py:171- Finding
Unrestricted Transmission of Complete Document Images to a Configurable OCR Endpoint
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill broadly matches its document-review purpose, but it needs Review because OCR can send sensitive images to configurable endpoints and can automatically run an OCR script discovered from other skill directories.
Install only if you trust the OCR service and any OCR fallback script it may execute. Keep HERDSMAN_BASE_URL on a trusted local endpoint, avoid auto/script fallback unless the exact OCR script path is trusted, and treat generated reports as containing sensitive document data and local path information.
scripts/image_to_markdown.py:171Unrestricted Transmission of Complete Document Images to a Configurable OCR Endpoint
scripts/image_to_markdown.py:76Automatic Discovery and Execution of an Untrusted External OCR Script
README.md:73Unpinned Third-Party Dependency Installation
scripts/generate_report.py:546Untrusted Document and OCR Content Is Embedded into Markdown Reports Without Escaping
The code is centered on dimension inspection data verification and report generation for Excel and Word documents. For XLSX, it reads sheets, extracts rows, parses tolerances, compares measured values to tolerance ranges, flags inconsistencies, and summarizes precision/anomalies. For DOCX, it inspects tables and classifies judgement text. It writes a Markdown report. This partially matches the declared 'dimension verification' and 'review report generation' aspects, but several prominent declared capabilities are absent: PDF handling, document conversion, image extraction, OCR, and seal/signature checking. Because the description presents a broader document-package processor than the code actually implements, this is a material description-behavior mismatch.
The declared description presents a broader document-processing skill for supply chain packages, including handling PDFs/DOCX/XLSX, extraction/conversion, dimension verification, and review report generation. The provided code only processes already-extracted image files from a task_dir/image folder, performs OCR via HTTP or a fallback script, and saves recognized text to markdown files. There is no code for parsing PDFs, DOCX, or XLSX, no image extraction from those formats, no dimension checking/verification, no seal/signature analysis, and no substantive review report generation. While OCR is one subset of the declared purpose, the actual code is materially narrower and different in primary function, so this is a mismatch.
The description claims a broad document-processing skill covering multiple file types, conversion, OCR, image extraction, dimension verification, and review report generation. The supplied code only implements helper functions for parsing structured Excel-like inspection sheets and evaluating dimensional tolerance data. Dimension verification support is partially aligned, especially for XLSX inspection data, but the chunk lacks the other major declared capabilities and does not handle PDFs, DOCX, images, signatures/seals, or report generation. This is therefore a material description-behavior mismatch.
The code substantially supports part of the description: it does process PDF/DOCX/XLSX-family documents, converts PDFs to DOCX, extracts images, and builds output artifacts for review. However, several key declared functions are absent. There is no dimension analysis/verification logic, no report-generation logic beyond an image index spreadsheet, no seal/signature detection, and no OCR engine invocation. For scanned PDFs, the script embeds page renderings into DOCX rather than extracting text via OCR. The code’s actual primary purpose is document conversion plus image extraction/indexing, not full dimension inspection or document review analysis. Therefore the declared description materially overstates the implemented behavior.
Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.
def script_ocr_image(image_path, ocr_script, config):
if not ocr_script:
raise RuntimeError("No fallback OCR script found")
proc_env = os.environ.copy()
proc_env["HERDSMAN_BASE_URL"] = normalize_base_url(config["base_url"])
proc = subprocess.run(
[
The README documents OCR HTTP calls and fallback behavior without explicitly warning that extracted image or document-derived data may be transmitted to an OCR service. In this skill's context—processing supply chain packages, signatures, seals, and inspection records—that omission is material because users may reasonably assume all processing is local, leading to unintended disclosure of sensitive business documents.
The README explicitly documents OCR over HTTP (POST /v1/ocr) but does not clearly constrain that traffic to a strictly local-only trust boundary or explain the data handling implications. Because this skill processes potentially sensitive supplier documents and extracted images, undocumented network transmission expands the attack surface and can expose confidential content if the endpoint is remote, misconfigured, or intercepted.
The skill instructs the agent to use shell, file I/O, environment variables, and networked OCR services, but it does not declare any explicit tool scope or allowed-tools boundary. That creates an over-privileged execution surface where an orchestrator may permit broader capabilities than users expect, increasing the risk of unintended file access, command execution, or data transfer.
The trigger phrases are broad enough to activate the skill for generic 'review', 'table data review', or document-related requests outside the intended scope. Over-broad activation can cause the agent to run file-processing, OCR, or shell-based workflows on unrelated user data, leading to unnecessary data exposure or unintended actions.
The OCR workflow states that images are sent to an HTTP service, but the skill does not clearly warn users that document-derived image data may leave the local processing context. Because supply-chain documents can contain proprietary designs, signatures, seals, and personal data, undisclosed transmission to a service materially increases confidentiality and compliance risk.
The OCR service URL is configurable via config and environment variables, so image contents may be sent to an arbitrary endpoint rather than a fixed local service. In a supply-chain document review skill, those images can contain proprietary drawings, signatures, and sensitive manufacturing data, making silent exfiltration a meaningful risk.
The script searches broad filesystem locations and an environment-supplied directory for scripts/ocr.py, then later executes the first match. That behavior exceeds normal document processing and turns OCR fallback into a code-loading mechanism, so any malicious or trojanized skill directory in those search roots can hijack execution.
The function uploads full image files to the configured OCR HTTP endpoint without any explicit consent or warning at the point of transmission. Given this skill processes supplier document packages, the data may include confidential engineering drawings or signatures, so undisclosed network transfer increases privacy and compliance risk.
The fallback mechanism executes an external ocr.py with little transparency to the caller about where that code came from or what privileges it has. Because the script path can originate from broad search paths or environment configuration, the lack of safety controls and disclosure materially increases the chance of running attacker-controlled code during ordinary OCR processing.
The subprocess call itself is not shell-injection prone because it uses an argument list, but it executes a Python script whose path is discovered dynamically from other skill directories or environment-controlled configuration. In this skill context, that means processing an image can trigger execution of untrusted code outside the current skill boundary, creating a realistic arbitrary-code-execution path if an attacker can place or point to a malicious ocr.py.
raise RuntimeError("No fallback OCR script found")
proc_env = os.environ.copy()
proc_env["HERDSMAN_BASE_URL"] = normalize_base_url(config["base_url"])
proc = subprocess.run(
[
"uv",
"run",
The generated _ImageIndex.xlsx includes absolute filesystem paths via img.resolve(), which can disclose host-specific directory layouts, usernames, project structure, or mounted storage locations to downstream users or systems that consume the report. That information is not required for the stated document-processing function, so it creates unnecessary sensitive data exposure and increases the blast radius if outputs are shared externally.
This markdown file explains that the skill creates output/, image/, imagetomd/, and ReviewReport.md, and later mentions archive copies of documents. However, it does not clearly warn users up front that running the skill will write multiple derived files into the provided task directory and may consume storage or alter the folder structure.
The exclusion rules explicitly include Chinese directory-name fragments 副本 and 复制, which encodes a locale-specific behavior in the skill documentation. The file does not indicate that this is optional, user-selectable, or justified as a region-specific tool, so it can be read as forcing a particular locale assumption.
This code computes a default report path and then opens it in write mode, which will create or overwrite the report file. Although this is part of the script's functionality, there is no explicit warning or confirmation that running the tool will write a file to disk at the resolved location.
This code creates output/ and image/ directories, copies source files, saves converted DOCX files, extracts embedded images, and generates an Excel index. Although the script prints file locations and processing status, it does not clearly warn the user up front that it will create and populate these directories under the provided task folder.
The script recursively processes all supported files under the provided task directory using rglob('*'), which can sweep in unrelated or sensitive documents beyond the user's intended subset. In a supply-chain document review context, broad collection increases the chance of over-processing confidential files and producing derived outputs for documents that were not meant to be handled.
Detected: suspicious.install_untrusted_source