Back to skill

Security audit

PDF和图片文字提取

Security checks for vulnerabilities and agentic risk

Overview

This skill does perform PDF and image text extraction, but it also requires a RedFox API key and remote permission check before local processing and understates some file-writing behavior.

Review before installing. Use it only if you are comfortable configuring a RedFox API key, allowing a network permission check on each use, and having document pages or images handled by the agent's vision tooling. Avoid confidential documents unless that external processing and local OCR-image/output-file creation are acceptable, and run it in a constrained workspace with pinned dependencies if possible.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:34
Finding

Mandatory External Authorization and User-Response Hijacking

Content
View full analysis
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
scripts/record.py:26
Finding

Unnecessary Transmission of an Environment Credential and Usage Metadata

Content
View full analysis
str: key = os.getenv("REDFOX_API_KEY", "").strip() if not key: print("This skill requires a RedFox API key for usage permission.") print(f"Registration address: {REGISTER_URL}") print("Set the REDFOX_API_KEY environment variable after obtaining a key.") sys.exit(1) return key def save_record() -> None: api_key = _get_api_key() payload = { "skillName": SKILL_NAME, "source": "PDF and Image Text Extraction-ClawHub", } headers = { "Content-Type": "application/json; charset=utf-8", "X-API-Key": api_key, } resp = requests.post( RECORD_URL, json=payload, headers=headers, verify=True, timeout=10, ) ``` The displayed string literals are translated into English; the control flow, endpoint, environment variable, request headers, and transmitted fields match the audited source. ### Technical Analysis The script reads a reusable API credential from the process environment and transmits it in the `X-API-Key` header to an external service. It also submits skill-identifying and source-identifying usage metadata. TLS certificate verification and a finite timeout are correctly enabled. No hidden upload of document contents was found. Nevertheless, this remote credential use is mandatory under the skill instructions even though PDF parsing, table extraction, scanned-page rendering, and batch processing are implemented locally. This violates least-privilege principles by requiring access to an environ ...[truncated 1401 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
SKILL.md:43
Finding

Unpinned Third-Party Dependencies and Unsafe Shell Installation Example

Content
View full analysis
=1.23.0 requests>=2.28.0 ``` ### Technical Analysis The installation command does not pin exact package versions and does not verify package hashes. It therefore resolves mutable future releases from the configured Python package index. A compromised, malicious, or unexpectedly incompatible future release could be installed without any corresponding change to the audited skill package. The version specifiers are also unquoted. In POSIX-compatible shells, the `>` characters can be interpreted as output-redirection operators rather than as part of the package requirement. This can alter the effective pip arguments and create local files such as `=1.23.0` or `=2.28.0`. As a result, the documented minimum versions may not be enforced as intended. No evidence of dependency confusion, typosquatting, or a currently malicious package was found. The confirmed issue is unsafe and non-reproducible dependency installation guidance. ### Attack Path 1. A user follows the documented installation command in a shell. 2. The shell may parse each `>` character as output redirection. 3. Pip receives package names without the intended version constraints, while redirection files may be created. 4. Alternatively, in an environment where the arguments remain intact, pip resolves any future release satisfying the broad lower bounds. 5. Package installation or import executes code supplied by the selected third-party distribution. 6. A compromised future release could consequently execute with the privileges of the user running pip or the Agent process. ### Impact Assessment A malicious dependency release could execute arbitrary Python or installation-time code with the privileges of the installing user. Depending on that user's ...[truncated 411 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (23)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The README promotes scanned-PDF recognition with AI vision and requires a third-party API key, but it does not clearly warn that document contents and images may be transmitted to an external service for processing. Because this skill handles potentially sensitive PDFs, contracts, reports, and research material, users could unknowingly exfiltrate confidential data to a remote provider.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill's stated purpose does not justify mandatory environment-variable access and outbound requests to a third-party service before local document processing can proceed. This creates an undisclosed data/control dependency on an external provider and enables account gating, telemetry, or future expansion of data exfiltration risk under the guise of a local utility.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill's stated purpose does not justify mandatory environment-variable access and outbound requests to a third-party service before local document processing can proceed. This creates an undisclosed data/control dependency on an external provider and enables account gating, telemetry, or future expansion of data exfiltration risk under the guise of a local utility.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill's stated purpose does not justify mandatory environment-variable access and outbound requests to a third-party service before local document processing can proceed. This creates an undisclosed data/control dependency on an external provider and enables account gating, telemetry, or future expansion of data exfiltration risk under the guise of a local utility.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The README says users can 'Simply describe what you need in natural language—no commands to memorize,' which does not clearly bound when the skill should activate or what phrasing is in scope. This broad trigger guidance increases the chance of unintended invocation from ordinary requests about PDFs or images.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The batch-processing phrase 'Extract text from all files in this folder' is overly generic and implies unconstrained directory-wide access. In a file-processing skill, this can lead to excessive data collection, accidental processing of sensitive files, or unintended traversal of large directories if the implementation does not enforce path and scope restrictions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The README says the skill can 'Save the extracted text' but does not explain where files are written, whether existing files may be overwritten, or whether user confirmation is required. In a local file-writing context, that omission can cause accidental overwrites, silent data creation in sensitive locations, or leakage of extracted content into insecure storage.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The README states '直接用自然语言描述需求即可,无需记忆命令', which does not clearly bound when this skill should activate versus when ordinary conversation should not. This kind of open-ended invocation guidance can overlap with common everyday speech and lacks explicit constraints or negative examples.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
80% confidence
Finding

The invocation examples and user instructions are presented only in Chinese and direct the user to interact in that language, but there is no explicit opt-in, alternative locale, or justification that the skill is intended only for a Chinese-language context. This can be a natural-language policy issue when a skill effectively forces a specific language without user choice.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The README mentions saving extracted results but does not warn that files may be created on disk or explain the destination path. This can lead to unintentional writes of sensitive OCR output, especially because extracted PDFs/images may contain confidential or regulated data, and unclear output locations increase the risk of accidental disclosure or overwrite.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill documents capabilities that require environment access, local file reads/writes, and network access, but it does not declare any explicit tool scope or allowed-tools boundary. That omission weakens reviewability and least-privilege enforcement, making it easier for the skill to exercise broader capabilities than users or hosts expect.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest frames the skill as a local PDF/image text extractor, yet the workflow requires remote authorization before any processing. That discrepancy is risky because operators may deploy the skill in restricted or privacy-sensitive environments assuming it works locally, when in fact it depends on a third-party service and may leak usage metadata.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

L019 将触发条件描述为“用户上传图片或 PDF 并要求提取文字,或询问文档中的文字内容;用户需要批量处理文件夹;用户需要提取 PDF 中的表格数据”,属于自然语言概述而非明确、封闭的触发短语集合。该描述缺少边界条件或排除示例,容易覆盖普通的文档问答、内容总结等相邻场景,增加技能被非预期调用的风险。

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Mandatory external authorization/account gating is not necessary for the stated document-extraction function and introduces avoidable privacy and availability risk. If the external service is unavailable or policy-disallowed, the skill becomes unusable; if metadata is transmitted, users may be exposed without realizing a local utility phones home.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The documentation claims processed files are not stored, but elsewhere instructs the skill to render scan pages to image files and optionally save extracted output to disk. This is a privacy and trust issue: sensitive document contents may persist locally contrary to user expectations, increasing exposure through leftover files, backups, or shared systems.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The script enforces a mandatory remote authorization and usage-recording call before the skill can be used, but the skill metadata describes OCR/PDF extraction functionality rather than a networked licensing or telemetry requirement. This creates an undisclosed data flow and remote dependency that can surprise users, leak usage metadata, and block local/offline use of an otherwise local processing skill.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The module docstring, usage guidance, and all runtime prompts are written only in Chinese, which imposes a specific language on users without opt-in. The policy explicitly disallows forced language or locale unless the skill offers a choice or clearly documents a justified region-specific constraint, which is not present here.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The code reads an API key from the environment and sends it to an external service for authentication even though the core stated purpose is document/image text extraction. This expands the trust boundary unnecessarily: secrets are exposed to a third party, usage becomes dependent on a remote service, and users may not expect credential handling in a local OCR utility.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
93% confidence
Finding

The script performs an external POST request carrying skill-identifying metadata and an API key header to a third-party endpoint. In the context of a document-extraction skill that users may expect to run locally, this network transmission is more dangerous because it introduces undisclosed telemetry, secret transmission, and a point of external control over tool availability.

Content

Scanner excerpt · scripts/record.py (reported line 52)May include surrounding context.

python
}

    try:
        resp = requests.post(RECORD_URL, json=payload, headers=headers, verify=True, timeout=10)
    except requests.exceptions.RequestException as e:
        print(f"❌ 无法连接鉴权服务器:{e}")
        print(f"请检查网络后重试,或前往 {REGISTER_URL} 确认 API Key 是否有效")

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

This file contains user-facing instructions and changelog text exclusively in Chinese, and the runtime output later in the script follows the same pattern. Under the policy, forcing a specific language without user opt-in is a natural-language policy concern unless the locale restriction is explicitly justified.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The manifest presents the skill as extracting text from images and PDFs, but this script persists state in the user's home directory via ~/.pdf_image_extractor_version. While minor, maintaining installation/use tracking state is outside the extraction behavior described in the manifest.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

save_version() writes to a file in the user's home directory to remember prior executions. Persisting per-user changelog state is not a direct or obvious requirement for extracting text from PDFs/images and is an ancillary capability beyond the stated purpose.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

When a page has little text, the script automatically creates an output directory and saves rendered page images to disk. Although this behavior is part of the implementation and exposed via CLI options, there is no clear user-facing warning in the code comments or interface text that running the tool may create image files on disk by default, which can matter for sensitive PDFs.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.