Back to skill

Security audit

arxiv文献解读,将提供arxiv链接,将文献翻译为中文解读,并给出对应的流程图

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent arXiv paper translation and diagram workflow, with expected network retrieval and output files but no hidden persistence, privilege escalation, or destructive behavior.

Install only if you are comfortable with the agent making outbound requests to paper and search services and saving generated Markdown/HTML files. Prefer using HTTPS arXiv API URLs, review generated HTML before sharing, and treat fetched paper content as untrusted source material.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:52
Finding

Unencrypted arXiv API Requests Permit Response Tampering

Content
View full analysis

Vulnerability Details

File Locations:

  • SKILL.md:52
  • references/arxiv-guide.md:13

Vulnerability Type: Use of an unencrypted HTTP connection for externally retrieved paper metadata
Risk Level: Medium

Vulnerable Code in SKILL.md:52:

bash
curl -s "http://export.arxiv.org/api/query?id_list=XXXXXXX" -H "User-Agent: Mozilla/5.0"

Vulnerable Code in references/arxiv-guide.md:13:

bash
curl -s "http://export.arxiv.org/api/query?id_list=XXXXXXX"

Technical Analysis

The documented workflow retrieves arXiv metadata and abstracts over plaintext HTTP. HTTP provides neither server authentication nor transport integrity. An attacker with a network position between the agent and the arXiv API—such as a malicious Wi-Fi operator, compromised proxy, or local network attacker—could intercept and modify the Atom response.

The workflow subsequently treats the retrieved title, author information, abstract, and related metadata as trusted input for translation, analysis, and diagram generation. Consequently, a modified response could cause the agent to produce falsified paper summaries or diagrams.

The retrieved response is treated as document content rather than executable code, so the reviewed instructions do not establish direct command execution or privilege escalation from this issue alone.

Attack Path

  1. A user supplies an arXiv identifier or URL.
  2. The agent follows the documented fallback procedure and requests the arXiv API endpoint over HTTP.
  3. A network-positioned attacker intercepts the plaintext request.
  4. The attacker returns a modified Atom response containing falsified metadata or abstract content.
  5. The agent translates and analyzes the altered content.
  6. The resulting Markdown report and HTML diagram present attacker-controlled information as though it came from arXiv.

Impact Assessment

The vulnerability primarily compromises the integrity and authe ...[truncated 419 chars]

Remediation
View remediation

Remediation Suggestions

  1. Replace every plaintext API URL with the HTTPS endpoint:

    bash
    curl --fail --silent --show-error \
      --proto '=https' \
      --max-redirs 0 \
      "https://export.arxiv.org/api/query?id_list=XXXXXXX" \
      -H "User-Agent: Mozilla/5.0"
    
  2. Update both SKILL.md and references/arxiv-guide.md so users and agents are not directed to the insecure endpoint.

  3. Fail closed on TLS certificate validation errors. Do not use options such as --insecure or -k.

  4. If redirects must be supported, restrict them to HTTPS destinations and an explicit allowlist of expected arXiv hosts.

  5. Validate the returned response format and ensure the requested arXiv identifier matches the identifier in the API response.

  6. Where output integrity is important, corroborate critical metadata through a second HTTPS-protected source, such as the arXiv abstract page or Crossref.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (8)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The description lists trigger phrases such as '论文结构图' and '生成论文图表' without clarifying when this skill should or should not activate. These phrases are broad enough to overlap with general requests about making paper diagrams, and the file provides no exclusion conditions or negative examples to constrain invocation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The description says the skill translates papers into Chinese and interprets arXiv papers as Chinese content, but it does not offer the user a language choice or indicate that Chinese output is optional. This is a natural-language locale policy issue because it forces a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill instructs the agent to retrieve content from external services, scrape web pages, and perform network searches without clearly warning the user that their request may trigger outbound requests. This can expose user-supplied paper identifiers, titles, or derived queries to third parties and may fetch untrusted remote content into the workflow.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

The Crossref API call is an external transmission path that sends request data to a third party. In this skill's context, outbound transmission is expected for paper retrieval, but it still creates privacy and trust-boundary risks because user-provided identifiers or derived metadata leave the local environment.

Content

Scanner excerpt · SKILL.md (reported line 65)May include surrounding context.

优先方案 A(Crossref API):

bash
curl -s "https://api.crossref.org/works/https://doi.org/10.48550/arXiv.XXXXXXX" -H "User-Agent: Mozilla/5.0"

方案 B(arXiv API):

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The entire guide is written in Chinese and does not indicate that other languages are supported or that Chinese is a required locale for a region-specific use case. Under the stated policy, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The document title and translation instructions are written as a Chinese-only workflow, including required Chinese translated headings and terminology mappings. This imposes a specific language/locale behavior without stating that the user can choose another language, which matches the policy category for language or locale violations.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · references/arxiv-guide.md (reported line 7)May include surrounding context.

1. Crossref API(优先)

bash
curl -s "https://api.crossref.org/works/https://doi.org/10.48550/arXiv.XXXXXXX"

返回:title, authors, abstract, DOI, date

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

The skill writes output artifacts, including HTML files, but does not disclose this behavior up front. Undisclosed file creation can surprise users, create persistence of sensitive derived content, and introduce downstream risk if generated HTML is later opened in a browser without sanitization controls.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.