Back to skill

Security audit

OCR Local V2

Security checks for vulnerabilities and agentic risk

Overview

This is a straightforward local OCR skill with disclosed first-run language-data downloads and no evidence of hidden data access, persistence, or destructive behavior.

Before installing, understand that OCR runs locally but first use may download Tesseract language data and cache it. For stricter reproducibility, prefer a version that pins tesseract.js and includes a lockfile or validated dependency version.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (7)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This markdown file presents all operational instructions, headings, and checklist content in Chinese, which effectively forces a specific language for users reading the skill publication guidance. The policy allows locale constraints only when the skill offers language choice or clearly documents a justified regional limitation, neither of which appears here.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The manifest describes the skill as '100% local, no API key required', which implies operation without remote service interaction. This file explicitly says the Tesseract language data is 'automatically downloaded at runtime', indicating network activity that contradicts the stronger local-only framing in the documentation context.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Lines L73 and L83 present conflicting intent: one claims 'pure local operation', while the first-run note states the user will download Tesseract language data on first use. That is an active contradiction in the documentation about whether network access is involved.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The manifest and README-style description repeatedly describe the skill as '100% local' and 'no API key required', implying no external service interaction. However, the notes explicitly state that the first run downloads language data, which means the skill performs network access at runtime and is not purely local in operation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The package description states that the skill supports Chinese (simplified/traditional) and English, which creates an explicit language scope in natural-language metadata without clarifying whether users can choose locale behavior or whether other languages are unsupported by design. This is a mild language/locale policy concern because the file presents a fixed language set rather than an opt-in or user-selectable language policy.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
92% confidence
Finding

Using a caret range for tesseract.js allows installation of different future minor/patch releases, which reduces build reproducibility and can unexpectedly introduce vulnerable or malicious dependency changes. In a local OCR skill, dependency integrity matters because the package processes untrusted files and executes third-party code in the runtime.

Content

Scanner excerpt · package.json (reported line 14)May include surrounding context.

json
"license": "MIT",
  "acceptLicenseTerms": true,
  "dependencies": {
    "tesseract.js": "^7.0.0"
  },
  "repository": {
    "type": "git",

Unverifiable Dependency: tesseract.js has 1 known advisory(ies) (GHSA-83rx-c8cr-6j8q (Insecure Default Configuration in tesseract.js)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
88% confidence
Finding

The manifest references tesseract.js without pinning an exact version, while static analysis notes a known advisory affecting some releases. Because the installed version is not fixed, consumers cannot verify whether they are protected, and the skill context increases concern because OCR libraries may parse complex attacker-supplied image data and rely on worker/runtime configuration.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.