Back to skill

Security audit

siliconflow-ocr

Security checks for vulnerabilities and agentic risk

Overview

This appears to be a real SiliconFlow OCR skill, but it needs Review because the documentation and shipped script disagree while user documents are sent to an external OCR service.

Review before installing if you plan to OCR private, regulated, or business-confidential documents. The skill sends image contents or image URLs to SiliconFlow, uses a SiliconFlow API key, and its documentation overstates or conflicts with the shipped script's actual model, PDF/batch support, and proxy handling.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (14)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

虽然代码与声明都围绕 SiliconFlow OCR,且都支持本地路径或 URL 作为输入,但存在多项实质性不一致。最关键的是主要实现路径与默认能力不同:声明强调默认走 PaddleOCR-VL-1.5,并支持图片/PDF、批量并发和坐标 token 清洗;实际代码只是调用通用 Qwen VL 聊天补全接口做单张图片文字提取。资源访问方式也与声明不符:没有 requests、没有 proxy、没有 secret-egress-proxy 注入逻辑。这些差异已超出“实现细节不同”的范围,属于对能力和默认行为的实质性误述。

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill performs network access, reads environment secrets, and documents writing OCR progress/results to disk, but it declares no explicit tool scope or permission boundaries. This creates a governance and least-privilege gap: users and runtimes cannot clearly constrain what the skill may access or exfiltrate.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill sends local file contents or remote-image-derived content to a third-party OCR API, but the user-facing description does not clearly warn that document data leaves the local/system boundary. This can expose sensitive PDFs, images, tables, or credentials embedded in documents to an external service without informed consent.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
94% confidence
Finding

This finding duplicates the same outbound request site and identifies the same external transmission risk: document contents are uploaded to a third-party API. The risk is contextual rather than overtly malicious, but remains security-relevant because the uploaded data may contain confidential information.

Content

Scanner excerpt · SKILL.md (reported line 48)May include surrounding context.

md
"temperature": 0.1,
        "max_tokens": 8000,
    }
    resp = requests.post(
        "https://api.siliconflow.cn/v1/chat/completions",
        json=payload,
        headers={"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"},

External Transmission

Medium
Category
Data Exfiltration
Confidence
94% confidence
Finding

This finding duplicates the same outbound request site and identifies the same external transmission risk: document contents are uploaded to a third-party API. The risk is contextual rather than overtly malicious, but remains security-relevant because the uploaded data may contain confidential information.

Content

Scanner excerpt · SKILL.md (reported line 48)May include surrounding context.

md
"temperature": 0.1,
        "max_tokens": 8000,
    }
    resp = requests.post(
        "https://api.siliconflow.cn/v1/chat/completions",
        json=payload,
        headers={"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"},

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

The hardcoded external endpoint confirms that OCR data is sent off-system to SiliconFlow. In an OCR skill this is functionally necessary, but it increases risk when handling sensitive documents because all recognized source material is processed by a remote provider under its own retention and logging policies.

Content

Scanner excerpt · SKILL.md (reported line 49)May include surrounding context.

md
"max_tokens": 8000,
    }
    resp = requests.post(
        "https://api.siliconflow.cn/v1/chat/completions",
        json=payload,
        headers={"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"},
        timeout=timeout,

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest claims broader functionality and a different default model than the code actually implements. In this file, the function accepts only one image path or one image URL, has no PDF handling, no batch/concurrency logic, no coordinate-token cleaning, and defaults to Qwen/Qwen3-VL-8B-Instruct rather than PaddleOCR-VL-1.5.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script sends local image contents or user-supplied image URLs to a third-party OCR service without any explicit consent flow, warning, or data-classification guard. In a skill context, users may pass sensitive documents and not realize the full content is being transmitted off-host, creating confidentiality and compliance risk.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

This skill is designed to exfiltrate image content to an external OCR API as part of normal operation, which is acceptable functionally but still a real data-transfer risk. Because the input can be a local file and the skill metadata mentions secret/proxy integration, the context increases sensitivity: users may run it in trusted environments and inadvertently send confidential material outside their boundary.

Content

Scanner excerpt · scripts/ocr.py (reported line 63)May include surrounding context.

python
data = json.dumps(payload).encode('utf-8')
    req = urllib.request.Request(
        'https://api.siliconflow.cn/v1/chat/completions',
        data=data,
        headers={
            'Content-Type': 'application/json',

Dynamic import via __import__()

Medium
Category
Dangerous Code Execution
Confidence
75% confidence
Finding

Dynamic import() can load arbitrary modules at runtime, bypassing static analysis and potentially importing malicious code.

Content

Scanner excerpt · scripts/ocr.py (reported line 127)May include surrounding context.

python
f.write(result['text'])
            print(f"\n已保存到: {args.output}")
    else:
        print(f"错误: {result['error']}", file=__import__('sys').stderr)
        exit(1)

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

The documented batch workflow says OCR progress/results may be incrementally written to progress.jsonl, but it does not warn users that recognized text and metadata can persist on local disk. This can create unintended data retention and secondary exposure of sensitive OCR output.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

The inline documentation explicitly states '使用 Qwen VL 模型进行图片文字识别', and the function docstring repeats that it uses SiliconFlow Qwen VL. This conflicts with the manifest's claim that the skill defaults to PaddleOCR-VL-1.5, creating a direct intent/documentation divergence about the model choice.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The default prompt explicitly instructs the model in Chinese to output all recognized text, and the CLI/tooling is presented only in Chinese. This creates a language/locale constraint without offering the user an explicit language choice or documenting that the skill is intentionally Chinese-only.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The command-line interface descriptions and help strings are entirely in Chinese, which effectively forces a specific locale for user interaction. Under the policy, this should either be optional, user-selectable, or clearly justified as a region-specific tool.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.