Back to skill

Security audit

paper-report

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent paper-report generator, but it needs review because its web downloads, dependency installs, and generated HTML outputs are under-scoped and can create security exposure.

Review this skill before installing. Use it only in a constrained workspace, avoid untrusted paper HTML pages, verify every figure URL before download, prefer pinned local dependencies, and treat generated HTML reports as untrusted until escaping and script controls are added.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/extract_figure_urls.py:58
Finding

Unrestricted Figure Retrieval Enables Server-Side Request Forgery

Content
View full analysis
]+src=["\']([^"\']+)["\']', block, re.IGNORECASE) img_urls = [] for src in srcs: if src.startswith('data:'): continue full_url = urljoin(base_url, src) img_urls.append(full_url) ``` The documented workflow subsequently retrieves the extracted URL: ```bash # reader/html.md:81-85 mkdir -p {workspace}/figures curl -sL "{FIGURE_URL}" -o {workspace}/figures/{FIGURE_NAME} ``` ### Technical Analysis Figure URLs originate in arbitrary HTML supplied through a user-selected webpage. `urljoin()` resolves absolute, scheme-relative, and relative URLs, but the result is not validated before it is presented for download. There are no restrictions on: - URL scheme - Destination hostname or port - Loopback, private, link-local, or reserved IP addresses - Cloud instance metadata addresses - Cross-origin image retrieval - Redirect destinations The documented `curl -L` operation follows redirects, so validating only the initial URL would also be insufficient. Although an agent is expected to select relevant figures manually, the workflow explicitly directs it to download URLs emitted by this parser and does not require a security review of those destinations. ### Attack Path 1. An attacker hosts an apparent academic paper or modifies a supported HTML paper page. 2. The page includes a `` containing an image such as: ```html Experimental results ``` 3. `extract_figure_urls.py` accepts and prints the URL without checking its destination. 4. Following `reader/html.md`, the agent invokes `curl -sL` on the extracted UR ...[truncated 992 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/crop_figures.py:49
Finding

Figure Name Path Traversal Allows Writes Outside the Output Directory

Content
View full analysis
= doc.page_count: print(f"WARNING: Page {page_idx} out of range (0-{doc.page_count - 1}), skipping '{name}'") continue page = doc[page_idx] rect = fitz.Rect(*rect_coords) # Validate rect is within page bounds page_rect = page.rect if not page_rect.contains(rect): print(f"WARNING: Rect {rect_coords} extends beyond page bounds {list(page_rect)}, clamping.") rect = rect & page_rect # intersection / clamp pix = page.get_pixmap(matrix=mat, clip=rect) out_path = os.path.join(output_dir, f"{name}.png") pix.save(out_path) ``` ### Technical Analysis The crop specification is parsed from a command-line JSON document, and each `name` value is used directly in an output path. `os.path.join()` does not enforce containment: - A name containing `../` can traverse to a parent directory. - An absolute name can cause the preceding output directory to be discarded. - Nested separators can target arbitrary subdirectories where parent paths already exist. The script validates the page and rectangle but does not validate the output name or compare the resolved destination against the intended output directory. The resulting content is always a rendered PNG, so this is not a fully attacker-selected byte write. It is nevertheless an overwrite primitive for `.png` paths accessible to the current process. ### Attack Path 1. A malicious or compromised workflow causes the crop specification to include a name such as: ```json { "page": 0, "rect": [0, 0, 100, 100], "name": "../../shared/important" } ``` 2. The script con ...[truncated 1003 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
writer/html-template.html:197
Finding

Unescaped Paper-Derived Content Can Produce Active HTML

Content
View full analysis
{{PAPER_TITLE}} — 阅读报告 ``` ```html

{{PAPER_TITLE_CN}}

{{PAPER_TITLE_EN}}

作者:{{AUTHORS}}
机构:{{AFFILIATIONS}}
来源:{{VENUE}} · {{YEAR}}

一、研究背景与动机

{{BACKGROUND_CONTENT}}

二、核心方法 / 技术方案

{{METHOD_CONTENT}}

Figure 1: 系统架构 图 1:{{FIG1_CAPTION}} ``` The writer documentation directs exact string replacement: ```text The template contains a complete , placeholders such as {{PAPER_TITLE_CN}}, {{FIG1_BASE64}}, and {{FIG1_CAPTION}}, and a section skeleton. Replace the placeholders using string replacement. ``` ### Technical Analysis Metadata, captions, and report text originate from a paper or webpage that may be attacker-controlled. The generation instructions use raw string substitution and provide no contextual encoding or HTML sanitization requirement. If a value contains markup such as `

`, the browser interprets it as active document content rather than text. Injection is possible in body, title, caption, and metadata placeholders. The validation checklist verifies document structur ...[truncated 1711 chars]

Remediation
View remediation
``` If MathJax is embedded inline, use hashes or nonces rather than broadly allowing arbitrary inline scripts. 6. Extend validation to fail on unexpected `

T08 · Insecure Dependencies

Warning
Location
writer/html-template.html:17
Finding

Unpinned Runtime Dependencies and Unverified Remote Browser Script

Content
View full analysis
``` ### Technical Analysis The Python and Node dependencies are installed without exact version constraints, integrity hashes, or a project lockfile. Consequently, identical Skill instructions may install different code over time. The `npm install -g` fallback also modifies the global Node environment rather than an isolated project environment. The MathJax URL includes a version but has no Subresource Integrity attribute. It remains an external runtime dependency: opening the generated report causes the browser to retrieve and execute JavaScript from the CDN. This also conflicts with the writer's description of the output as self-contained. There is no evidence that the named dependencies are intentionally malicious. The vulnerability is the mutable, insufficiently verified supply-chain boundary. ### Attack Path Package-installation path: 1. The required dependency is absent. 2. The agent follows the documented `pip install` or `npm install -g` instruction. 3. The package registry supplies the current package release and its transitive dependencies without project-enforced hashes. 4. A compromised package, account, ...[truncated 966 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (20)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

While most TP4 instances look like false positives, this variant correctly notes an undeclared capability issue: the skill directs local .docx writing and local image/file reads without any corresponding permission declaration. Misdescribed capabilities can mislead operators about what the skill will access or modify, which matters in environments that rely on manifest metadata for trust decisions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

While most TP4 instances look like false positives, this variant correctly notes an undeclared capability issue: the skill directs local .docx writing and local image/file reads without any corresponding permission declaration. Misdescribed capabilities can mislead operators about what the skill will access or modify, which matters in environments that rely on manifest metadata for trust decisions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

While most TP4 instances look like false positives, this variant correctly notes an undeclared capability issue: the skill directs local .docx writing and local image/file reads without any corresponding permission declaration. Misdescribed capabilities can mislead operators about what the skill will access or modify, which matters in environments that rely on manifest metadata for trust decisions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

While most TP4 instances look like false positives, this variant correctly notes an undeclared capability issue: the skill directs local .docx writing and local image/file reads without any corresponding permission declaration. Misdescribed capabilities can mislead operators about what the skill will access or modify, which matters in environments that rely on manifest metadata for trust decisions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

While most TP4 instances look like false positives, this variant correctly notes an undeclared capability issue: the skill directs local .docx writing and local image/file reads without any corresponding permission declaration. Misdescribed capabilities can mislead operators about what the skill will access or modify, which matters in environments that rely on manifest metadata for trust decisions.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 98)May include surrounding context.

md
- **Word (.docx) 输出** → [writer/docx.md](writer/docx.md)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill explicitly instructs use of network access (curl to arXiv) and local file read/write operations for inputs, figures, and report outputs, but it declares no tool scope or permission boundaries. Missing capability declarations increase the risk of over-privileged or opaque execution, making it harder to constrain what the skill may access or modify.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The invocation description is broad enough to trigger on generic requests like 'summarize' or 'analyze,' which can cause the skill to activate outside its intended paper-report niche. Over-broad triggering can unexpectedly route user content into file/network-processing workflows, increasing the chance of unintended data handling and user confusion.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The description explicitly says the skill converts papers into structured Chinese reading reports, and later instructions require the full report to be in Chinese. This forces a specific language choice without indicating that the user can opt in, opt out, or select another language.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The writing requirement says "全文中文", which is a hard language constraint. Because no opt-in or alternative language path is provided in this file, it conflicts with the policy requirement to avoid forcing a locale or language without user choice.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

L077 明确写明“用中文记录摘要”,这是对输出语言的强制要求。根据规则,未经用户选择或正当区域性说明而限定语言,属于自然语言层面的语言/locale 政策问题。

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The code invokes execSync('npm root -g') to discover a global module path. Even though the command string is constant, spawning a shell introduces avoidable command-execution behavior and trusts environment-dependent resolution (PATH, npm shim, shell), which can be hijacked in compromised or multi-tenant runtimes to execute an attacker-controlled npm binary. In a paper-report skill, this is unnecessary and broadens the attack surface beyond document generation.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The instruction to ensure the output does not contain text such as 'AI-assisted generation' directs the agent to suppress provenance/disclosure language. This can mislead reviewers about how the document was produced, weaken auditability, and facilitate policy evasion in environments that require transparency for AI-generated content.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The HTML template declares lang="zh-CN" and uses Chinese-only titles, labels, and footer text throughout the document. This forces a specific language/locale by default with no visible opt-in, fallback, or indication that the skill is intentionally region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

L001 明确规定输出格式说明为中文语境下的“输出格式:HTML”,且全文未提供语言可选项或用户 opt-in 机制。根据规则,强制特定语言而没有选择或合理限制说明,属于自然语言层面的语言/地区策略违规。

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

该 markdown 文件适用 SQP-1。L010 将“用户未指定输出格式时”设为默认选择,但没有说明何时应优先 HTML、何时不应使用该格式,也没有给出排除条件,属于模糊激活条件,可能与很多普通写作请求重叠。

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The template hard-codes Chinese section headers and metadata labels such as '作者', '机构', and '研究背景与动机', which effectively requires generated content to be presented in Chinese. This is a natural-language locale constraint, and the file does not offer a user opt-in or alternative language choice, nor does it document a justified region-specific requirement.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

This markdown file instructs the skill to fetch remote HTML content using curl, which is a network operation covered by the missing-warning rule for markdown files. The document explains how to perform the download but does not include any warning that the skill will contact external URLs and retrieve remote content into the workspace.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

Step H3 tells the user to create a figures directory and download image files from extracted URLs, which affects local files and performs additional network access. The markdown describes the mechanics but omits a clear disclosure that the process will write multiple files and contact external hosts for figure retrieval.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dangerous_exec

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
writer/docx-template.js:27