Back to skill

Security audit

AutoDimension Report Skill En

Security checks for vulnerabilities and agentic risk

Overview

The skill broadly matches its document-review purpose, but it needs Review because OCR can send sensitive images to configurable endpoints and can automatically run an OCR script discovered from other skill directories.

Install only if you trust the OCR service and any OCR fallback script it may execute. Keep HERDSMAN_BASE_URL on a trusted local endpoint, avoid auto/script fallback unless the exact OCR script path is trusted, and treat generated reports as containing sensitive document data and local path information.

Vulnerability Patterns
  • Tool Hijacking and SpoofingModifies or replaces tools so legitimate-looking calls execute attacker logic
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/image_to_markdown.py:171
Finding

Unrestricted Transmission of Complete Document Images to a Configurable OCR Endpoint

Content
View full analysis
Remediation
View remediation

T07 · Tool Hijacking and Spoofing

Error
Location
scripts/image_to_markdown.py:76
Finding

Automatic Discovery and Execution of an Untrusted External OCR Script

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
README.md:73
Finding

Unpinned Third-Party Dependency Installation

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/generate_report.py:546
Finding

Untrusted Document and OCR Content Is Embedded into Markdown Reports Without Escaping

Content
View full analysis
{4}".format( issue["severity"].upper(), issue["type"], issue.get("sheet", ""), issue.get("item", "")[:24], issue["message"], ) ) ``` ```python if data["ocr_hits"]: w("**{0}** potential seal/signature keyword hits detected:".format(len(data["ocr_hits"]))) for hit in data["ocr_hits"][:20]: preview = " | snippet: `{0}`".format(hit["preview"]) if hit["preview"] else "" w("- `{0}` | hits: `{1}`{2}".format(hit["path"], ", ".join(hit["keywords"]), preview)) ``` The verification report similarly writes unescaped spreadsheet data into Markdown tables: ```python for row in file_rows[:200]: cells = [ str(row["seq"]), row["sheet"], row["item"][:20], row["std"][:18], ...[truncated 2983 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (21)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The code is centered on dimension inspection data verification and report generation for Excel and Word documents. For XLSX, it reads sheets, extracts rows, parses tolerances, compares measured values to tolerance ranges, flags inconsistencies, and summarizes precision/anomalies. For DOCX, it inspects tables and classifies judgement text. It writes a Markdown report. This partially matches the declared 'dimension verification' and 'review report generation' aspects, but several prominent declared capabilities are absent: PDF handling, document conversion, image extraction, OCR, and seal/signature checking. Because the description presents a broader document-package processor than the code actually implements, this is a material description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents a broader document-processing skill for supply chain packages, including handling PDFs/DOCX/XLSX, extraction/conversion, dimension verification, and review report generation. The provided code only processes already-extracted image files from a task_dir/image folder, performs OCR via HTTP or a fallback script, and saves recognized text to markdown files. There is no code for parsing PDFs, DOCX, or XLSX, no image extraction from those formats, no dimension checking/verification, no seal/signature analysis, and no substantive review report generation. While OCR is one subset of the declared purpose, the actual code is materially narrower and different in primary function, so this is a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The description claims a broad document-processing skill covering multiple file types, conversion, OCR, image extraction, dimension verification, and review report generation. The supplied code only implements helper functions for parsing structured Excel-like inspection sheets and evaluating dimensional tolerance data. Dimension verification support is partially aligned, especially for XLSX inspection data, but the chunk lacks the other major declared capabilities and does not handle PDFs, DOCX, images, signatures/seals, or report generation. This is therefore a material description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
91% confidence
Finding

The code substantially supports part of the description: it does process PDF/DOCX/XLSX-family documents, converts PDFs to DOCX, extracts images, and builds output artifacts for review. However, several key declared functions are absent. There is no dimension analysis/verification logic, no report-generation logic beyond an image index spreadsheet, no seal/signature detection, and no OCR engine invocation. For scanned PDFs, the script embeds page renderings into DOCX rather than extracting text via OCR. The code’s actual primary purpose is document conversion plus image extraction/indexing, not full dimension inspection or document review analysis. Therefore the declared description materially overstates the implemented behavior.

Content

No source excerpt is available for this finding.

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · scripts/image_to_markdown.py (reported line 212)May include surrounding context.

python
def script_ocr_image(image_path, ocr_script, config):
    if not ocr_script:
        raise RuntimeError("No fallback OCR script found")
    proc_env = os.environ.copy()
    proc_env["HERDSMAN_BASE_URL"] = normalize_base_url(config["base_url"])
    proc = subprocess.run(
        [

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The README documents OCR HTTP calls and fallback behavior without explicitly warning that extracted image or document-derived data may be transmitted to an OCR service. In this skill's context—processing supply chain packages, signatures, seals, and inspection records—that omission is material because users may reasonably assume all processing is local, leading to unintended disclosure of sensitive business documents.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

The README explicitly documents OCR over HTTP (POST /v1/ocr) but does not clearly constrain that traffic to a strictly local-only trust boundary or explain the data handling implications. Because this skill processes potentially sensitive supplier documents and extracted images, undocumented network transmission expands the attack surface and can expose confidential content if the endpoint is remote, misconfigured, or intercepted.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill instructs the agent to use shell, file I/O, environment variables, and networked OCR services, but it does not declare any explicit tool scope or allowed-tools boundary. That creates an over-privileged execution surface where an orchestrator may permit broader capabilities than users expect, increasing the risk of unintended file access, command execution, or data transfer.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The trigger phrases are broad enough to activate the skill for generic 'review', 'table data review', or document-related requests outside the intended scope. Over-broad activation can cause the agent to run file-processing, OCR, or shell-based workflows on unrelated user data, leading to unnecessary data exposure or unintended actions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The OCR workflow states that images are sent to an HTTP service, but the skill does not clearly warn users that document-derived image data may leave the local processing context. Because supply-chain documents can contain proprietary designs, signatures, seals, and personal data, undisclosed transmission to a service materially increases confidentiality and compliance risk.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The OCR service URL is configurable via config and environment variables, so image contents may be sent to an arbitrary endpoint rather than a fixed local service. In a supply-chain document review skill, those images can contain proprietary drawings, signatures, and sensitive manufacturing data, making silent exfiltration a meaningful risk.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The script searches broad filesystem locations and an environment-supplied directory for scripts/ocr.py, then later executes the first match. That behavior exceeds normal document processing and turns OCR fallback into a code-loading mechanism, so any malicious or trojanized skill directory in those search roots can hijack execution.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The function uploads full image files to the configured OCR HTTP endpoint without any explicit consent or warning at the point of transmission. Given this skill processes supplier document packages, the data may include confidential engineering drawings or signatures, so undisclosed network transfer increases privacy and compliance risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The fallback mechanism executes an external ocr.py with little transparency to the caller about where that code came from or what privileges it has. Because the script path can originate from broad search paths or environment configuration, the lack of safety controls and disclosure materially increases the chance of running attacker-controlled code during ordinary OCR processing.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
94% confidence
Finding

The subprocess call itself is not shell-injection prone because it uses an argument list, but it executes a Python script whose path is discovered dynamically from other skill directories or environment-controlled configuration. In this skill context, that means processing an image can trigger execution of untrusted code outside the current skill boundary, creating a realistic arbitrary-code-execution path if an attacker can place or point to a malicious ocr.py.

Content

Scanner excerpt · scripts/image_to_markdown.py (reported line 214)May include surrounding context.

python
raise RuntimeError("No fallback OCR script found")
    proc_env = os.environ.copy()
    proc_env["HERDSMAN_BASE_URL"] = normalize_base_url(config["base_url"])
    proc = subprocess.run(
        [
            "uv",
            "run",

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The generated _ImageIndex.xlsx includes absolute filesystem paths via img.resolve(), which can disclose host-specific directory layouts, usernames, project structure, or mounted storage locations to downstream users or systems that consume the report. That information is not required for the stated document-processing function, so it creates unnecessary sensitive data exposure and increases the blast radius if outputs are shared externally.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

This markdown file explains that the skill creates output/, image/, imagetomd/, and ReviewReport.md, and later mentions archive copies of documents. However, it does not clearly warn users up front that running the skill will write multiple derived files into the provided task directory and may consume storage or alter the folder structure.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The exclusion rules explicitly include Chinese directory-name fragments 副本 and 复制, which encodes a locale-specific behavior in the skill documentation. The file does not indicate that this is optional, user-selectable, or justified as a region-specific tool, so it can be read as forcing a particular locale assumption.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

This code computes a default report path and then opens it in write mode, which will create or overwrite the report file. Although this is part of the script's functionality, there is no explicit warning or confirmation that running the tool will write a file to disk at the resolved location.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

This code creates output/ and image/ directories, copies source files, saves converted DOCX files, extracts embedded images, and generates an Excel index. Although the script prints file locations and processing status, it does not clearly warn the user up front that it will create and populate these directories under the provided task folder.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The script recursively processes all supported files under the provided task directory using rglob('*'), which can sweep in unrelated or sensitive documents beyond the user's intended subset. In a supply-chain document review context, broad collection increases the chance of over-processing confidential files and producing derived outputs for documents that were not meant to be handled.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.install_untrusted_source

Install source points to URL shortener or raw IP.

Warn
Code
suspicious.install_untrusted_source
Location
scripts/config.json:2