Back to skill

Security audit

图片文字识别

Security checks for vulnerabilities and agentic risk

Overview

The skill is a real local OCR tool, but its “fully offline” promise is contradicted by default automatic network downloads and local caching of extracted text.

Review before installing if you require strict offline handling. Use --no-auto-download for OCR runs, avoid processing sensitive documents unless you are comfortable with local OCR text being cached on disk, and prefer preinstalled or hash-verified language models instead of runtime downloads.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (21)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill's stated scope is narrower than its documented behavior: it can fetch models, infer language from system locale/history, persist preferences, and emit structured metadata beyond plain text extraction. Such undeclared capability expansion increases privacy and integrity risk because users may expose documents under false assumptions about locality and data handling.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill's stated scope is narrower than its documented behavior: it can fetch models, infer language from system locale/history, persist preferences, and emit structured metadata beyond plain text extraction. Such undeclared capability expansion increases privacy and integrity risk because users may expose documents under false assumptions about locality and data handling.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

Claiming the skill is 'fully offline' while it may download language models over the network is materially misleading. Users may send sensitive scans or internal documents assuming no network egress, creating confidentiality and supply-chain exposure when remote fetches occur.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 62)May include surrounding context.

md
- `scripts/ocr_image.js` — OCR command line tool

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill metadata promises fully offline OCR, but this script fetches language models from GitHub at runtime. That discrepancy creates a real security and privacy risk: using the skill may unexpectedly transmit network metadata and introduces a remote supply-chain dependency for code-adjacent model assets.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill advertises fully offline local OCR, but its own CLI exposes --download-langs and --auto-download, enabling network retrieval of language models. This is dangerous because users may provide sensitive documents under an offline-only trust assumption, while the tool can initiate unexpected network activity and violate data-handling or compliance requirements.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The code automatically downloads models based on system language and prior usage history, which directly contradicts the skill's offline-only description. In a skill intended for OCR of potentially sensitive images and PDFs, hidden network behavior materially increases privacy, compliance, and trust risk even if document contents are not directly uploaded.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill declares no explicit tool/permission scope even though its documented behavior requires shell access and reads environment/system settings. That weakens containment and transparency, making it easier for the skill to execute broader host actions than users or reviewers expect.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill description emphasizes offline OCR but does not prominently warn that language-model downloads may occur over the network. This omission undermines informed consent and can cause users to process sensitive images under inaccurate assumptions about data isolation and egress.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

Automatic downloading of language models introduces unnecessary network and supply-chain risk for a skill presented as text-extraction OCR. Because the download is automatic, it can occur without meaningful user awareness, potentially exposing usage context and pulling unpinned remote artifacts into a trusted workflow.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Selecting language models based on system locale and previously used languages without opt-in leaks environment-derived information into skill behavior and can trigger unexpected downloads or persistence. In a document-processing context, silent adaptation based on host settings reduces predictability and may expose user preferences or regional metadata.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The inline Chinese description explicitly includes screenshots ("截图"), which conflicts with the higher-level policy that this skill should not be used for screenshots when text is directly selectable/copied. This inconsistency can cause misrouting and over-invocation of OCR on content that should be handled by safer or more appropriate text extraction paths, increasing the chance of bypassing intended workflow constraints.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The default prompt is broad and generic, asking to read text and numbers from an image without encoding the intended restrictions on when OCR should be used. This can make activation boundaries unclear and lead the agent to invoke the skill for general image-reading tasks or screenshots outside its approved scope.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

A network download path is inconsistent with a local/offline OCR skill and expands the attack surface beyond local processing. Even though the URL is fixed to the official tessdata_fast repository, the script still enables external fetches, exposing users to availability failures, privacy leakage, and upstream tampering risk if the remote content changes or the transport chain is compromised.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The only natural-language comment in the file is written in Chinese, which can violate language/locale policy when the skill does not offer language choice or document a justified locale constraint. This may make the script less understandable to users or maintainers who expect the default project language.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The user-facing usage and option descriptions are entirely in Chinese, and the script does not provide an opt-in or switch for another interface language. This imposes a locale/language choice on users rather than offering one, which matches the policy's language-choice concern.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The skill inspects system locale and persisted language preferences to influence later behavior, including model selection and potential downloads. While this is not code execution by itself, it collects and persists environment-derived metadata unrelated to core text extraction, which creates avoidable privacy leakage and makes the hidden network behavior more adaptive.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill invokes bash to run download_langs.sh, and elsewhere spawns platform tools like defaults, powershell, pdftoppm, and sips. While PDF rendering and preprocessing can be implementation details, shelling out to manage downloads and inspect system settings exceeds the manifest's simple 'text extraction only, fully offline' framing.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script's user-facing comments and status/error messages are written in Chinese, including operational output such as download progress and failure notices. This imposes a specific language on users without opt-in or justification, which matches the natural-language locale policy violation criteria.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The script invokes a subprocess via Node to run ocr_image.js, which is a safety-relevant operation for code files under this rule. There is no confirmation prompt, user-facing log message, or explanatory comment disclosing that the script will execute another program; the nearby comment only discusses stderr filtering in Chinese, not user disclosure.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The file's primary natural-language description is entirely in Chinese and does not indicate that other languages are supported or that the locale restriction is intentional. Under the policy criteria, forcing a specific language without user opt-in can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dangerous_exec

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
scripts/ocr_image.js:211