Back to skill

Security audit

多论文创新点对比

Security checks for vulnerabilities and agentic risk

Overview

This paper-comparison skill is mostly purpose-aligned, but it asks agents to run commands without confirmation and includes a self-evolution workflow that can persist changes into future skill behavior.

Review before installing. Only use it on a chosen paper folder, expect it to create progress and output files there, and do not allow automatic package installation or skill self-modification without explicit human review. Prefer running it in an isolated environment and disable or ignore the diary/PR self-evolution workflow.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:19
Finding

Unpinned Third-Party Dependency Installation

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, line 19
Vulnerability Type: Unverified runtime dependency installation
Risk Level: Medium

Evidence:

text
1. Confirm that python-docx is installed: `pip install python-docx`

The snippet above is an English rendering of the instruction at the specified source location.

Technical Analysis

The Skill instructs the agent to install python-docx directly from the configured Python package index. It does not specify a reviewed version, require package hashes, use a lockfile, constrain the package source, or require an isolated environment.

Because package installation can execute package build and installation logic, the effective code being trusted is not fully determined by the audited project. A compromised package release, compromised package index, malicious index configuration, or unexpected future dependency change could introduce code that was not present during this audit.

The package name is consistent with the expected legitimate library; no evidence of typosquatting or an intentionally malicious package was found. The risk arises from the unsafe, unpinned installation procedure.

Attack Path

  1. A user invokes the Skill on a system where python-docx is unavailable.
  2. Following SKILL.md, the agent executes pip install python-docx.
  3. pip resolves the package and its transitive dependencies from the environment's configured package index.
  4. If the index, selected release, dependency chain, or local package-index configuration is compromised, malicious installation or package code is retrieved.
  5. That code executes with the privileges of the account running the agent or is later executed when imported.

Exploitation therefore depends on compromise or manipulation of the relevant Python supply chain or package-index configuration.

Impact Assessment

Successful exploitation could execute arbitrary code with the privi ...[truncated 501 chars]

Remediation
View remediation

Remediation Suggestions

  1. Pin python-docx and all transitive dependencies to reviewed versions.
  2. Maintain a lockfile or requirements file containing cryptographic hashes, and install with hash verification.
  3. Use an explicitly trusted package index rather than inheriting an unknown environment configuration.
  4. Install dependencies inside a dedicated virtual environment with the minimum required permissions.
  5. Do not install packages automatically. Inform the user of the required dependency and obtain explicit approval before changing the environment.
  6. Prefer a prebuilt, reviewed execution environment in which dependencies are already installed.
  7. Periodically scan pinned packages for known vulnerabilities and update them through a controlled review process.

T02 · Agent Memory Poisoning

Warning
Location
SKILL.md:195
Finding

Persistent Self-Modification Can Promote Untrusted Content into Future Skill Instructions

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 195-200
Vulnerability Type: Persistent instruction and memory poisoning
Risk Level: Medium

Evidence:

text
## Self-evolution mechanism
After each execution of this Skill:
1. Evaluate whether the output achieved its goal: pass or fail.
2. On failure, append the failure case and remediation suggestion to diary/YYYY-MM-DD.md.
3. When a remediation suggestion recurs during the latest three executions, promote it into a formal rule and submit a pull request modifying this SKILL.md.

The snippet above is an English rendering of the instructions at the specified source location.

Technical Analysis

The Skill processes papers that must be treated as untrusted input. Its self-evolution mechanism directs the agent to persist model-generated failure analysis and remediation advice, then potentially promote repeated advice into the Skill's formal instructions.

This creates a trust-boundary violation between untrusted document content and durable agent instructions. A crafted paper could influence the agent's interpretation of a failure and the remediation advice written to the diary. Repeating the same influence across several executions could satisfy the stated promotion criterion.

The packaged Python script does not implement the diary or pull-request workflow, and the instructions only require submission of a pull request rather than automatic merging. Exploitation therefore depends on an agent following the textual workflow and on the proposed change being accepted or otherwise applied. Nevertheless, the Skill does not require sanitization, provenance tracking, security review, or explicit human approval before generated advice is proposed as a formal rule.

Attack Path

  1. An attacker supplies one or more crafted papers containing content designed to influence the agent's analysis or operational behavior.
  2. Processing the crafted content c ...[truncated 1443 chars]
Remediation
View remediation

Remediation Suggestions

  1. Remove automatic or recurrence-based promotion of model-generated observations into SKILL.md.
  2. Keep execution diaries separate from executable Skill instructions and mark all diary entries as untrusted analytical data.
  3. Require explicit, informed human approval before creating any proposed change to Skill instructions.
  4. Require security review, provenance records, and a clear diff for every proposed instruction change.
  5. Prevent source-document content from being copied into rules without sanitization and independent validation.
  6. Use an allowlisted schema for observations, excluding commands, tool directives, external URLs, and instruction-like language.
  7. Require signed commits and branch protection so an agent cannot directly alter the trusted Skill definition.
  8. Add tests ensuring that paper content cannot change agent policies, tool permissions, or future-session instructions.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (10)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill explicitly instructs execution of Bash commands and Python scripts without user confirmation, despite modifying files. This removes an important consent checkpoint and increases the risk of unintended filesystem changes, dependency installation, command misuse, or script-based abuse if paths, filenames, or supporting scripts are manipulated.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill performs file reads and writes, invokes Bash, and runs Python, but declares no explicit tool scope or permissions boundary. In practice this weakens least-privilege controls and makes it easier for the skill to access or modify user files beyond what is necessary, especially because it is designed to persist artifacts into user-controlled folders.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill directs the agent to write progress files and generated reports into the user's paper folder without warning or confirmation. This is dangerous because it can alter user workspaces, overwrite existing files, leak derived content into synced/shared directories, and normalize silent side effects during what appears to be a read-oriented analysis task.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The self-evolution mechanism instructs the skill to append diary entries and potentially modify its own SKILL.md via PRs, which is unrelated to the user task of comparing papers. This creates unauthorized persistence and opens a path for prompt/data contamination, where sensitive user content or flawed model judgments become embedded into future behavior or repository state.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The default prompt is phrased as a generic invocation directive and does not encode clear boundaries such as required input type, folder-size enforcement, or explicit task constraints. In an agent system, overly broad routing text can cause the skill to be invoked outside its intended context, increasing the chance of inappropriate file access, unintended bulk processing, or prompt-scope confusion.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This code writes the generated summary template with Chinese section headings and labels, imposing a specific language on users. The file provides no option to select another language and no documented reason that the tool must be Chinese-only.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The README states that the skill will automatically write summ.md, paper_summ_state.json, and extracted_text/ into the user-provided paper folder, but it does not prominently warn that running the skill modifies that directory. This can lead to unexpected file creation or overwriting in user data locations, especially when the folder contains important research artifacts or is under version control.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The manifest presents the skill as an analysis/comparison tool, but the implementation also performs persistent writes such as progress tracking and report generation in the user's folder. While these writes are functionally related, the mismatch can surprise users and create integrity/privacy issues if they do not expect new files to be created beside their source papers.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

The instruction '使用中文输出:报告和对话均使用中文' forces a specific language for both report and conversation. This is a natural-language policy concern because it does not provide user opt-in or an alternative locale choice.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
75% confidence
Finding

The display name and description are in Chinese, while the default prompt is written in English. This creates an implicit language behavior inconsistency without documenting user choice, opt-in, or a justified locale policy, which can conflict with language preference expectations.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.