Back to skill

Security audit

Paper Summarize Pdf To Feishu

Security checks for vulnerabilities and agentic risk

Overview

The skill’s PDF-to-Feishu workflow is mostly coherent, but it asks for elevated installs and exposes paper contents and Feishu document tokens in ways that need careful review.

Install only if you are comfortable with a Chinese-language workflow that creates Feishu documents, uploads extracted paper content/images, stores local copies and logs, and may print document tokens. Run it in a controlled workspace with dependencies installed ahead of time rather than following inline sudo commands, and avoid using it on confidential or attacker-supplied PDFs until prompt-injection boundaries, token redaction, path validation, and JSON escaping are fixed.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (4)

T01 · Skill Instruction Hijacking

Warning
Location
SKILL.md:769
Finding

Forced and Potentially False Identity Attribution in Generated Documents

Content
View full analysis
| | 生成时间 | | | 文档版本 | v1.0(最终版) | **审核修改记录**: - ✅ 核实接受日期:已确认 - ✅ 关键数据已核对(X 处一致,Y 处已补充) - ✅ 图片已上传(X 张) ``` ``` ### Technical Analysis The skill explicitly requires the agent to append a fixed author and model identity—`Lux (qwencode/qwen3.5-plus)`—to every final document. This attribution is not derived from the active agent, configured model, authenticated user, or actual document author. Because the instruction is mandatory and executed during finalization, it modifies the agent's output independently of the user's request and can produce a false provenance claim. The separate template requirement in `references/summary_template.md:190-207` reinforces the requirement that an attribution block must always be present. This is instruction-level output hijacking rather than local code execution. It does not grant operating-system privileges, but it compromises the integrity and provenance of every document generated through the skill. ### Attack Path 1. A user invokes the skill to summarize a PDF. 2. The skill completes extraction, summarization, review, and Feishu document creation. 3. During the finalization stage, the controlling agent follows the mandatory attribution instruction. 4. The agent appends `Lux (qwencode/qwen3.5-plus)` even when that identity and model did not create the document. 5. Recipients are presented with inaccurate authorship and model provenance. ### Impact Assessment - **Privileges obtained**: No additional system privileges. - **Affected scope ...[truncated 361 chars]
Remediation
View remediation

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:197
Finding

Prompt Injection Through Untrusted PDF Content Passed to Tool-Capable Sub-Agents

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/extract_metadata.sh:76
Finding

JSON Injection Through Unescaped PDF Metadata

Content
View full analysis
"$OUTPUT_JSON" << EOF { "paper_id": "$PAPER_ID", "doi": "$DOI", "title": "$TITLE", "authors": "$AUTHOR", "subject": "$SUBJECT", "keywords": "$KEYWORDS", "producer": "$PRODUCER", "created_date": "$CREATEDATE", "pages": $PAGES, "pdf_hash": "$PDF_HASH", "source_file": "$INPUT_PDF", "extracted_at": "$(date -Iseconds)" } EOF ``` ### Technical Analysis Values extracted from the PDF and its filename are inserted into a JSON document through a shell here-document without JSON escaping. The affected values include title, author, subject, keywords, producer, creation date, DOI, paper ID, and source path. PDF metadata is attacker-controlled. A value containing a double quote, backslash, newline, or JSON syntax can terminate the intended string and inject additional properties. At minimum, this creates malformed JSON and denies downstream processing. Depending on duplicate-key handling and injected field placement, an attacker may also influence values consumed later by `jq`, including `paper_id`, `doi`, or `source_file`. Shell quoting does not solve JSON encoding. The script must use a JSON serializer that correctly escapes every string and validates numeric fields. ### Attack Path 1. An attacker creates a PDF with a crafted metadata value, such as a title containing quotes and additional JSON properties. 2. The victim runs `extract_metadata.sh` on the PDF. 3. `pdfinfo` returns the malicious metadata. 4. The script inserts the value verbatim into `metadata.json`. 5. The output becomes malformed or contains attacker-shaped properties. 6. `check_duplicate.sh` subsequently consumes the file with `jq`. 7. Downstream behavior is disrupted or manipulated according to the injected values and parser behavior. ...[truncated 568 chars]
Remediation
View remediation
"$OUTPUT_JSON" ``` 2. Validate that `PAGES` contains digits before passing it through `--argjson`. 3. Validate the completed document with `jq empty "$OUTPUT_JSON"` before allowing downstream use. 4. Enforce a schema and reject duplicate or unexpected properties. 5. Limit metadata field lengths and remove prohibited control characters where appropriate. 6. Treat every value returned by `pdfinfo` and `pdftotext` as attacker-controlled. ]]>

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/check_duplicate.sh:23
Finding

Path Traversal Through Unvalidated Metadata Paper Identifier

Content
View full analysis
Remediation
View remediation
&2 exit 1 fi ``` 2. Explicitly reject `/`, backslashes, absolute paths, null-like values, and `..` path components. 3. Canonicalize both paths and verify containment: ```bash PAPERS_ROOT=$(realpath -e "$PAPERS_DIR") PAPER_DIR=$(realpath -m "$PAPERS_ROOT/$PAPER_ID") case "$PAPER_DIR/" in "$PAPERS_ROOT"/*) ;; *) echo "Path escapes papers directory" >&2; exit 1 ;; esac ``` 4. Consider symbolic links: use `realpath -e` after the target exists and re-check containment before reading files. 5. Validate the metadata file against a schema before extracting any path-related values. 6. Do not print complete Feishu tokens. Redact them in logs and output unless disclosure is strictly required. 7. Open sensitive files only after validation, preferably using a trusted directory file descriptor or equivalent containment-safe mechanism. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (21)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的是一个面向最终用户的完整论文处理与总结工作流,而实际代码仅实现其中非常底层的一步:PDF 元数据提取与标识生成。虽然 DOI、哈希和 paper_id 可能作为后续去重/处理的支撑信息,但该片段本身并不执行总结、写入飞书文档、图表处理、审核确认等核心能力。因此代码行为与声明的主要目的存在明显不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

声明描述的是一个较完整的论文 PDF 到飞书文档总结系统,包含长流程编排、去重、总结、配图、审核与人工确认等能力。而实际代码仅处理“补充材料”PDF:复制文件、提取文本、简单 grep 关键词、截取前100行生成 markdown 摘要,并提示用户之后再用 feishu_doc 追加到文档。它没有直接操作飞书文档,也没有实现完整论文总结或审核能力。因此该代码片段的实际行为只覆盖声明中的很小一部分辅助环节,且其主要目的更偏向补充材料预处理,和声明的核心能力存在明显不匹配。

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill instructs use of sudo apt-get install to modify the host system and install packages with elevated privileges. A document-processing skill should not require inline privilege escalation, because this materially increases blast radius if the skill is invoked on a sensitive machine or if package sources or commands are altered.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
95% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · scripts/locate_figures.sh (reported line 135)May include surrounding context.

sh
fi

# 清理临时文件
rm -f "$OUTPUT_DIR"/embedded_temp-*.png 2>/dev/null || true

echo ""
echo "✅ Figure 定位完成"

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill creates persistent local directories and logs containing processing artifacts, but the description does not disclose this storage behavior. While less severe than external upload, it still affects privacy and retention expectations and can leave sensitive paper contents on disk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill description, trigger phrases, workflow prompts, placeholders, and required output templates are all specified in Chinese, and the instructions mandate Chinese-specific formatting such as using Chinese brackets. There is no user opt-in or documented locale restriction justifying a Chinese-only experience.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrases are broad enough to match ordinary PDF-related conversations, increasing the chance the skill activates when the user did not intend a Feishu-uploading, multi-step workflow. In context, accidental activation matters because the skill creates files, logs content locally, and may upload extracted text and images externally.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
97% confidence
Finding

This line explicitly directs sudo execution for package installation, introducing privileged system modification into a document summarization skill. Elevated execution is dangerous because it can alter the host environment, expand compromise impact, and normalize unsafe trust in skill-authored commands.

Content

Scanner excerpt · SKILL.md (reported line 84)May include surrounding context.

bash
# PDF 处理工具
sudo apt-get install -y poppler-utils

# OCR 工具
sudo apt-get install -y tesseract-ocr tesseract-ocr-eng tesseract-ocr-chi-sim

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
97% confidence
Finding

This is another explicit sudo-based install command, again requiring elevated privileges unrelated to the narrow user task of summarizing a paper. Repeated privileged instructions increase the chance an operator executes them reflexively, magnifying host-level risk.

Content

Scanner excerpt · SKILL.md (reported line 87)May include surrounding context.

md
sudo apt-get install -y poppler-utils

# OCR 工具
sudo apt-get install -y tesseract-ocr tesseract-ocr-eng tesseract-ocr-chi-sim

# JSON 处理工具(脚本中大量使用)
sudo apt-get install -y jq

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
97% confidence
Finding

This command continues the pattern of instructing root-level package installation from inside the skill workflow. Even benign package names are problematic here because the skill should not prompt privilege escalation as part of ordinary content-processing execution.

Content

Scanner excerpt · SKILL.md (reported line 90)May include surrounding context.

sudo apt-get install -y tesseract-ocr tesseract-ocr-eng tesseract-ocr-chi-sim

JSON 处理工具(脚本中大量使用)

sudo apt-get install -y jq

text

**快速检查**:

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The skill reads ~/.openclaw/openclaw.json from the user's home directory to auto-select a model, which reaches beyond the immediate PDF summarization input and inspects local user configuration. Even if read-only, this expands data access scope and may expose environment details, provider names, or model settings to the skill workflow unnecessarily.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The Reviewer sub-agent prompt directly injects raw slices of the source paper and the generated summary into another model context. This broad duplication of content increases exposure of potentially sensitive or unpublished material and is especially risky when the fallback model or provider may differ from the primary processing path.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

Stage four instructs sending the audit report contents back to the user wholesale, which can amplify prior over-collection by redisclosing detailed excerpts or sensitive findings unnecessarily. If the paper contains confidential data or reviewer prompts include copied source text, this broadens the leak surface further.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The template explicitly requires the summary title to use a Chinese title placeholder, which signals a fixed language requirement. Because the file does not offer a language option or justify a locale-specific constraint, this is a natural-language policy violation under the language/locale rule.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

This shell script uses Chinese exclusively in its comments, usage text, and operational output, which is a natural-language locale constraint. Under the policy, forcing a specific language without user opt-in or documented justification is a violation, and no alternative language or opt-in mechanism is provided here.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This shell script’s natural-language comments and console/log messages are written in Chinese, including the usage text and status output. The file does not indicate that the skill is region-specific or provide any user opt-in or language choice, which conflicts with the policy against forcing a specific language without user consent.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This shell script's comments, usage text, log messages, and warnings are entirely in Chinese, which imposes a specific language on users without any opt-in or documented locale requirement. The policy explicitly flags language or locale constraints when the skill does not offer a choice or justify the restriction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script hard-codes Tesseract to use English (-l eng) and only validates OCR output against English figure labels (Figure/Fig.). This imposes a specific language assumption in behavior rather than offering a language choice or documenting a justified locale restriction.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The script prints the Feishu document token directly to stdout in a user-facing 'next steps' message. In automation environments, stdout is commonly captured by CI logs, agent traces, terminal history, or centralized logging systems, which can unintentionally expose the token to other users or systems and enable unauthorized document access or modification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill creates persistent local directories and logs containing processing artifacts, but the description does not disclose this storage behavior. While less severe than external upload, it still affects privacy and retention expectations and can leave sensitive paper contents on disk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

Natural-language policy violations apply to all file types, including comments and output strings in code. This script consistently uses Chinese for usage text, errors, status updates, and next-step instructions, with no opt-in, fallback, or documentation that the skill is intentionally region-specific.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.