Back to skill

Security audit

pdf-ocr

Security checks for vulnerabilities and agentic risk

Overview

This OCR skill is mostly purpose-aligned, but it can silently upload local documents to a cloud OCR service after local OCR failure and can install Python packages at runtime.

Review before installing, especially if you handle contracts, IDs, financial records, medical records, or confidential business documents. Use it only in an isolated environment, preinstall reviewed dependencies yourself, avoid configuring a SiliconFlow API key unless you intend cloud OCR, and do not rely on the default RapidOCR path as guaranteed offline because the code can fall back to cloud processing.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/pdf_ocr_processor.py:153
Finding

Local OCR failure silently falls back to cloud processing and transmits document contents

Content
View full analysis
Dict[str, Any]: """Use the SiliconFlow API to recognize a PDF.""" images = self.pdf_to_images(pdf_path) text_parts = [] for idx, img_base64 in enumerate(images, 1): page_text = self.siliconflow_engine.recognize(img_base64, idx) text_parts.append(f"=== Page {idx} ===\n{page_text}") return { "text": "\n\n".join(text_parts), "page_count": len(images), "engine": "silico ...[truncated 3597 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
scripts/pdf_ocr_processor.py:22
Finding

Missing dependencies are installed automatically at runtime without version or hash pinning

Content
View full analysis
=1.23.0 pillow>=9.0.0 requests>=2.28.0 python-dotenv>=1.0.0 rapidocr_onnxruntime>=1.3.0 ``` ### Technical Analysis The Skill invokes `python -m pip install ` automatically when imports fail. The installed packages are identified only by package name, without an exact version, integrity hash, locked transitive dependency set, or explicitly constrained package index. As a resul ...[truncated 2021 chars]
Remediation
View remediation
=version` in production installation manifests. 9. Run the Skill as a non-administrative account with minimal filesystem and network access so that a compromised dependency has limited reach. 10. Replace runtime installation with an error such as: ```python try: from rapidocr_onnxruntime import RapidOCR except ImportError as exc: raise RuntimeError( "rapidocr_onnxruntime is required for local OCR. " "Install dependencies from the project's locked environment file." ) from exc ``` ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (52)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · .gitignore (reported line 40)May include surrounding context.

text
Thumbs.db

# Environment
.env
.env.local

# Testing

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · .gitignore (reported line 41)May include surrounding context.

text
# Environment
.env
.env.local

# Testing
.pytest_cache/

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The documentation introduces a cloud OCR engine but does not clearly warn that document images/text will be sent to a third-party service for processing. Users may unknowingly expose sensitive PDFs, contracts, IDs, or other confidential material, creating a significant privacy and compliance risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill states it will automatically switch from the local OCR engine to the cloud API when RapidOCR initialization fails, without describing any explicit user warning or consent flow. This creates a dangerous silent fallback where sensitive local documents may be exfiltrated remotely merely because a local dependency is missing or broken.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

This OCR skill can alter the host environment by installing Python packages on demand, a capability not necessary for ordinary document recognition. Runtime installation can pull and execute unreviewed third-party code, create nondeterministic behavior, and bypass deployment controls expected by users of a data-processing skill.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

When local RapidOCR initialization fails, the processor silently switches to the cloud OCR engine and may transmit document contents to a remote service. This is dangerous because users selecting a local engine likely expect data to remain on-host; the fallback changes the trust boundary without explicit consent and can leak sensitive PDFs.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The README promotes a cloud OCR mode that sends document content to a third-party API, but it does not clearly warn users that uploaded PDFs/images may contain sensitive data and will leave the local environment. In an OCR skill, this matters because users may process contracts, IDs, scans, or other confidential documents and assume behavior is equivalent to the local engine.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The documented trigger phrases for selecting the local engine are broad natural-language expressions such as requests for quick/offline/local processing. In an agent skill context, overly broad routing phrases can cause unintended engine selection behavior, making execution dependent on ambiguous wording rather than explicit user confirmation or structured parameters.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The cloud-engine trigger phrases are ambiguous and can overlap with ordinary requests like 'high accuracy' or 'process complex scans,' which could silently route documents to a remote OCR provider. In this skill's context, that is more dangerous than a normal misrouting bug because unintended cloud selection can expose sensitive document contents to third-party processing and possible billing.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill documentation describes capabilities that can access environment variables, local files, network resources, and shell execution, but it does not declare any explicit tool scope or allowed-tools boundary. In an agent environment, this increases the chance of over-privileged execution and makes it harder to constrain what the skill may do when invoked.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The trigger phrases for the local engine are broad everyday requests such as '快速识别这个文档' and '离线处理这个 PDF', which can overlap with normal user language. This can cause unintended skill activation, leading the agent to access files or begin OCR processing when the user did not specifically mean to invoke this skill.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The cloud-engine trigger phrases are even more generic, including requests like '高精度识别这个文档' and '用云端 OCR 引擎', which may match many ordinary conversations. Because this path can transmit document contents to a third-party API, unintended activation creates a stronger privacy and data-exfiltration risk than a purely local engine.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill is presented as an OCR extraction tool, but the implementation also changes the execution environment by installing dependencies automatically. This mismatch reduces user awareness of privileged behavior and makes the skill more dangerous in context because operators may grant it access expecting passive document processing only.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
96% confidence
Finding

The code invokes pip at runtime via subprocess to install packages automatically. Even though the package names are hardcoded in current call sites, this still gives the skill unexpected environment-modification capability and executes network-fetched code during normal OCR operation, which materially expands the attack surface and violates least privilege for an OCR utility.

Content

Scanner excerpt · scripts/pdf_ocr_processor.py (reported line 27)May include surrounding context.

python
"""自动安装缺失的依赖"""
    print(f"正在安装依赖: {package}")
    try:
        subprocess.check_call([sys.executable, "-m", "pip", "install", package])
        print(f"依赖 {package} 安装成功")
        return True
    except subprocess.CalledProcessError as e:

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/pdf_ocr_processor.py (reported line 83)May include surrounding context.

python
def __init__(self, api_key: str = "", model: str = "deepseek-ai/DeepSeek-OCR"):
        self.api_key = api_key or os.getenv("SILICON_FLOW_API_KEY", "")
        self.model = model or os.getenv("SILICON_FLOW_OCR_MODEL", "deepseek-ai/DeepSeek-OCR")
        self.base_url = "https://api.siliconflow.cn/v1/chat/completions"
        self.headers = {
            "Content-Type": "application/json",
            "Authorization": f"Bearer {self.api_key}"

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The cloud OCR path sends base64-encoded document images to an external API without a clear user-facing warning or consent at the point of transmission. OCR inputs often contain sensitive material, so undisclosed remote transfer creates confidentiality and compliance risks, especially in enterprise or regulated environments.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
94% confidence
Finding

This request transmits OCR input data to a third-party endpoint. External transmission is expected for a cloud OCR feature, but in this skill it is security-relevant because document contents may be sensitive and the upload can occur after silent fallback from local mode, making the data exposure less obvious to the user.

Content

Scanner excerpt · scripts/pdf_ocr_processor.py (reported line 126)May include surrounding context.

python
}
        
        try:
            response = requests.post(
                self.base_url, 
                headers=self.headers, 
                json=payload,

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

The guide repeatedly instructs users to run npx skills ... without pinning the package or version being executed. This creates a supply-chain risk because npx may fetch whatever version is current at execution time, allowing unexpected behavior or malicious package compromise to affect installers.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

This installation example invokes npx skills without a pinned version, so users may execute a different package revision than intended. If the upstream package is hijacked or updated with unsafe behavior, the command could run attacker-controlled code during install.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

Using npx skills with a full URL still leaves the executed CLI itself unpinned. That means the trust boundary includes whatever version npx resolves at runtime, which can expose users to package substitution or compromised updates.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

This command narrows the installed skill but still launches an unpinned npx skills package. The risk remains that the installer binary changes over time or is maliciously replaced, leading to arbitrary code execution in the user's environment.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
95% confidence
Finding

The global installation command is more sensitive because it may modify a broader environment scope while still relying on an unpinned npx skills package. A compromised upstream release could therefore affect all users or agents using the global install.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
95% confidence
Finding

The non-interactive --yes example amplifies the danger of an unpinned npx execution because it suppresses user confirmation while fetching a potentially changing package. In CI/CD, this can make supply-chain compromise automatic and widespread.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
93% confidence
Finding

This direct install example again depends on an unpinned npx skills package, exposing users to registry-side or upstream compromise. Because the command is presented as a primary install path, it meaningfully increases operational risk.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
92% confidence
Finding

Even the search command runs an unpinned npx skills package, which still executes code from an unfixed upstream version. While lower impact than installation, it can still expose users to unexpected code execution locally.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.