Back to skill

Security audit

Paper Reader Deep

Security checks for vulnerabilities and agentic risk

Overview

This is a disclosed Chinese-language PDF paper-reading helper that creates local Markdown reports and may query Crossref for DOI metadata, with no evidence of hidden, destructive, credential-seeking, or deceptive behavior.

Install only if you are comfortable with a Chinese-language workflow that reads PDFs from a directory, writes report files alongside them, and may disclose DOI values to Crossref. For sensitive unpublished papers, avoid networked DOI lookup or run in a controlled/offline environment; for stricter environments, pin dependencies in a virtual environment before use.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
README.md:16
Finding

Unpinned Third-Party Dependencies

Content
View full analysis

Vulnerability Details

File Location: README.md:16-20
Vulnerability Type: Unpinned dependency installation
Risk Level: Medium

Vulnerable Code:

bash
## Installation Dependencies

```bash
pip install pdfplumber PyYAML
text

### Technical Analysis

The documented installation command retrieves mutable package versions from the user's configured Python package index. Neither dependency is constrained to a reviewed version, and the project provides no lock file or cryptographic hashes for artifact verification.

As a result, the code installed by users can differ from the code that was reviewed during this audit. If a dependency publisher account, package repository, package release, or local package-index configuration is compromised, an attacker-controlled distribution could be selected during dependency resolution. Python packages may execute code during installation or when imported by `scripts/deep_reader.py`.

This finding concerns supply-chain integrity. The audit found no evidence that `pdfplumber` or `PyYAML` is currently malicious.

### Attack Path

1. An attacker compromises a dependency publisher account, package repository, upstream release process, or package index used by the victim.
2. The attacker publishes or serves a malicious version of `pdfplumber`, `PyYAML`, or one of their transitive dependencies.
3. A user follows the project documentation and runs `pip install pdfplumber PyYAML`.
4. Pip resolves and downloads the attacker-controlled artifact because no reviewed versions or hashes are enforced.
5. Malicious code executes during package installation or when the dependency is imported and used.
6. The payload runs with the permissions of the user performing the installation or invoking the skill.

### Impact Assessment

Successful exploitation could provide arbitrary code execution under the installing user's account. The resulting scope may include access to files readable
...[truncated 426 chars]
Remediation
View remediation

Remediation Suggestions

  1. Create a dependency file that pins every direct and transitive dependency to a reviewed version.
  2. Generate and record SHA-256 hashes for all approved distributions.
  3. Require hash verification during installation:
    bash
    python -m pip install --require-hashes -r requirements.txt
    
  4. Use a lock-file generation tool such as pip-tools and review dependency changes before updating the lock file.
  5. Prefer an approved package index and explicitly configure trusted repository sources in controlled environments.
  6. Run dependency vulnerability and provenance checks in CI.
  7. Install dependencies inside an isolated virtual environment without administrative privileges.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (14)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The title and the entire instruction set are written in Chinese, and the output filenames are also specified in Chinese, with no indication that users may choose another language. This can violate language or locale policy when the skill imposes a specific language without opt-in.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The activation condition is overly broad: any user request to deeply read or generate a paper report could trigger the skill without clear scoping or confirmation. In context, this can cause unintended processing of files or invocation in situations where the user only wanted discussion or a lightweight summary, increasing the chance of unauthorized file access or surprising side effects.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill explicitly writes per-PDF reports to the source directory and updates MEMORY.md, but it does not require prior user consent or prominently warn about filesystem modifications. This is dangerous because activating the skill could create or alter files in user directories unexpectedly, potentially leaking sensitive derived content, overwriting expected workspace state, or violating least-surprise expectations.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The entire skill is written as a Chinese-only analysis framework, including headings, prompts, and writing requirements, with no indication that the user can choose another language. This creates a language/locale policy concern because the skill implicitly mandates a specific language without opt-in or justification for a region-specific use case.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

This markdown template is entirely written in Chinese and includes a required field labeled '中文标题', which indicates the skill expects or enforces a specific language/locale. The file does not offer the user a language choice or explain that the template is intended only for a Chinese-language or region-specific context.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This YAML config defines the skill's report structure entirely with Chinese section names such as "基础信息", "科学问题", and "对你研究的意义". Because the file provides no indication that language is selectable or that the skill is intentionally limited to Chinese-language use, it appears to impose a specific language/locale by default, which matches the policy-violation category for forced language without user opt-in.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The module docstring presents the tool as a '论文深度阅读报告生成器' and '最终版', implying full report generation. However, the generated report explicitly fills major sections with placeholders like '[AI分析中...]' and later states that manual or further AI processing is still needed, which contradicts the stated capability rather than merely omitting detail.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This Python file contains natural-language strings, docstrings, status messages, usage text, and generated report content exclusively in Chinese. Because there is no option to select another language or document that the skill is region-specific, it imposes a locale/language choice on users, which matches the policy-violation criteria.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
96% confidence
Finding

The script extracts a DOI from local PDF text and sends it to Crossref over the network without explicit user consent or an offline-only option. While a DOI is usually low sensitivity, this still creates external data transmission from local content and can leak what documents a user is processing, which is more concerning in research, unpublished manuscript, or restricted-environment contexts.

Content

Scanner excerpt · scripts/deep_reader.py (reported line 65)May include surrounding context.

python
doi = doi.strip()
            if not doi.startswith('10.'):
                return None
            url = f"https://api.crossref.org/works/{doi}"
            req = urllib.request.Request(url, headers={'User-Agent': 'Mozilla/5.0'})
            with urllib.request.urlopen(req, timeout=5) as response:
                data = json.loads(response.read())

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file’s title and usage instruction are entirely in Chinese, and the template directs use in that language without offering any language or locale choice. This is a natural-language policy concern because it imposes a specific language by default rather than making language selection optional or explicitly justified as region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This markdown template presents all headings and instructions exclusively in Chinese and does not offer an alternative language or indicate that the skill is region-specific. Under the policy rule, forcing a specific language without user choice is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

This markdown file states that processing will generate multiple report files in the source PDF directory. Because this behavior modifies the user's filesystem, the skill description should explicitly disclose it as a side effect or caution so users understand files will be created alongside their PDFs.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

Line L110 makes a concrete behavioral claim about what will and will not be recorded in MEMORY.md. In the provided file, this appears as intent/documentation, but there is no implementation here to substantiate or enforce that boundary, so the documentation overstates a specific behavior constraint.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.