T01 · Skill Instruction Hijacking
- Location
SKILL.md:34- Finding
Mandatory External Authorization and User-Response Hijacking
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill does perform PDF and image text extraction, but it also requires a RedFox API key and remote permission check before local processing and understates some file-writing behavior.
Review before installing. Use it only if you are comfortable configuring a RedFox API key, allowing a network permission check on each use, and having document pages or images handled by the agent's vision tooling. Avoid confidential documents unless that external processing and local OCR-image/output-file creation are acceptable, and run it in a constrained workspace with pinned dependencies if possible.
SKILL.md:34Mandatory External Authorization and User-Response Hijacking
scripts/record.py:26Unnecessary Transmission of an Environment Credential and Usage Metadata
SKILL.md:43Unpinned Third-Party Dependencies and Unsafe Shell Installation Example
The README promotes scanned-PDF recognition with AI vision and requires a third-party API key, but it does not clearly warn that document contents and images may be transmitted to an external service for processing. Because this skill handles potentially sensitive PDFs, contracts, reports, and research material, users could unknowingly exfiltrate confidential data to a remote provider.
The skill's stated purpose does not justify mandatory environment-variable access and outbound requests to a third-party service before local document processing can proceed. This creates an undisclosed data/control dependency on an external provider and enables account gating, telemetry, or future expansion of data exfiltration risk under the guise of a local utility.
The skill's stated purpose does not justify mandatory environment-variable access and outbound requests to a third-party service before local document processing can proceed. This creates an undisclosed data/control dependency on an external provider and enables account gating, telemetry, or future expansion of data exfiltration risk under the guise of a local utility.
The skill's stated purpose does not justify mandatory environment-variable access and outbound requests to a third-party service before local document processing can proceed. This creates an undisclosed data/control dependency on an external provider and enables account gating, telemetry, or future expansion of data exfiltration risk under the guise of a local utility.
The README says users can 'Simply describe what you need in natural language—no commands to memorize,' which does not clearly bound when the skill should activate or what phrasing is in scope. This broad trigger guidance increases the chance of unintended invocation from ordinary requests about PDFs or images.
The batch-processing phrase 'Extract text from all files in this folder' is overly generic and implies unconstrained directory-wide access. In a file-processing skill, this can lead to excessive data collection, accidental processing of sensitive files, or unintended traversal of large directories if the implementation does not enforce path and scope restrictions.
The README says the skill can 'Save the extracted text' but does not explain where files are written, whether existing files may be overwritten, or whether user confirmation is required. In a local file-writing context, that omission can cause accidental overwrites, silent data creation in sensitive locations, or leakage of extracted content into insecure storage.
The README states '直接用自然语言描述需求即可,无需记忆命令', which does not clearly bound when this skill should activate versus when ordinary conversation should not. This kind of open-ended invocation guidance can overlap with common everyday speech and lacks explicit constraints or negative examples.
The invocation examples and user instructions are presented only in Chinese and direct the user to interact in that language, but there is no explicit opt-in, alternative locale, or justification that the skill is intended only for a Chinese-language context. This can be a natural-language policy issue when a skill effectively forces a specific language without user choice.
The README mentions saving extracted results but does not warn that files may be created on disk or explain the destination path. This can lead to unintentional writes of sensitive OCR output, especially because extracted PDFs/images may contain confidential or regulated data, and unclear output locations increase the risk of accidental disclosure or overwrite.
The skill documents capabilities that require environment access, local file reads/writes, and network access, but it does not declare any explicit tool scope or allowed-tools boundary. That omission weakens reviewability and least-privilege enforcement, making it easier for the skill to exercise broader capabilities than users or hosts expect.
The manifest frames the skill as a local PDF/image text extractor, yet the workflow requires remote authorization before any processing. That discrepancy is risky because operators may deploy the skill in restricted or privacy-sensitive environments assuming it works locally, when in fact it depends on a third-party service and may leak usage metadata.
L019 将触发条件描述为“用户上传图片或 PDF 并要求提取文字,或询问文档中的文字内容;用户需要批量处理文件夹;用户需要提取 PDF 中的表格数据”,属于自然语言概述而非明确、封闭的触发短语集合。该描述缺少边界条件或排除示例,容易覆盖普通的文档问答、内容总结等相邻场景,增加技能被非预期调用的风险。
Mandatory external authorization/account gating is not necessary for the stated document-extraction function and introduces avoidable privacy and availability risk. If the external service is unavailable or policy-disallowed, the skill becomes unusable; if metadata is transmitted, users may be exposed without realizing a local utility phones home.
The documentation claims processed files are not stored, but elsewhere instructs the skill to render scan pages to image files and optionally save extracted output to disk. This is a privacy and trust issue: sensitive document contents may persist locally contrary to user expectations, increasing exposure through leftover files, backups, or shared systems.
The script enforces a mandatory remote authorization and usage-recording call before the skill can be used, but the skill metadata describes OCR/PDF extraction functionality rather than a networked licensing or telemetry requirement. This creates an undisclosed data flow and remote dependency that can surprise users, leak usage metadata, and block local/offline use of an otherwise local processing skill.
The module docstring, usage guidance, and all runtime prompts are written only in Chinese, which imposes a specific language on users without opt-in. The policy explicitly disallows forced language or locale unless the skill offers a choice or clearly documents a justified region-specific constraint, which is not present here.
The code reads an API key from the environment and sends it to an external service for authentication even though the core stated purpose is document/image text extraction. This expands the trust boundary unnecessarily: secrets are exposed to a third party, usage becomes dependent on a remote service, and users may not expect credential handling in a local OCR utility.
The script performs an external POST request carrying skill-identifying metadata and an API key header to a third-party endpoint. In the context of a document-extraction skill that users may expect to run locally, this network transmission is more dangerous because it introduces undisclosed telemetry, secret transmission, and a point of external control over tool availability.
}
try:
resp = requests.post(RECORD_URL, json=payload, headers=headers, verify=True, timeout=10)
except requests.exceptions.RequestException as e:
print(f"❌ 无法连接鉴权服务器:{e}")
print(f"请检查网络后重试,或前往 {REGISTER_URL} 确认 API Key 是否有效")
This file contains user-facing instructions and changelog text exclusively in Chinese, and the runtime output later in the script follows the same pattern. Under the policy, forcing a specific language without user opt-in is a natural-language policy concern unless the locale restriction is explicitly justified.
The manifest presents the skill as extracting text from images and PDFs, but this script persists state in the user's home directory via ~/.pdf_image_extractor_version. While minor, maintaining installation/use tracking state is outside the extraction behavior described in the manifest.
save_version() writes to a file in the user's home directory to remember prior executions. Persisting per-user changelog state is not a direct or obvious requirement for extracting text from PDFs/images and is an ancillary capability beyond the stated purpose.
When a page has little text, the script automatically creates an output directory and saves rendered page images to disk. Although this behavior is part of the implementation and exposed via CLI options, there is no clear user-facing warning in the code comments or interface text that running the tool may create image files on disk by default, which can matter for sensitive PDFs.
No suspicious patterns detected.