Back to skill

Security audit

Skill Evolution

Security checks for vulnerabilities and agentic risk

Overview

This skill is a disclosed editor for other skills, but it can permanently add document- or conversation-derived rules into future agent behavior, so it needs careful review before use.

Install only if you intend to let this skill edit other installed skills. Before applying any proposed fusion or evolution, inspect the exact source text and final diff, avoid feeding it confidential documents or sensitive conversation content, and clean up .skill-backups logs if they may contain private material.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T02 · Agent Memory Poisoning

Warning
Location
SKILL.md:181
Finding

Persistent Instruction Injection Through Untrusted Skill and Document Ingestion

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 181-230
Vulnerability Type: Persistent poisoning of skill instructions
Risk Level: Medium

Vulnerable Instruction Excerpt

The following is an English translation of the relevant source instructions:

text
B2.2: Ingestion from an external document

1. The user provides an absolute document path.
2. Parse according to format: .md directly, .docx with python-docx,
   .pdf with pdfplumber/PyMuPDF, and .txt directly.
3. Extract:
   - Split by paragraph or section.
   - Identify tables, lists, and rules.
   - Evaluate the value of each block to the target skill.
4. Generate a candidate list containing:
   - A summary of the knowledge block.
   - Its source, including file name and location.
   - The dimension it may enhance after integration.

B2.3: Ingestion from conversation experience

1. Analyze user corrections to skill output in the current conversation.
2. Extract the user's preferred output format.
3. Extract professional knowledge supplied by the user.
4. Refine it into a reusable pattern.
5. Generate a candidate list.

B4.2: Integration

For each selected item:
1. Adapt the knowledge block to the style of the target skill.
2. Mark its source.
3. Insert it at the recommended location.
4. Handle conflicts:
   - If it contradicts existing logic, pause and ask the user to decide.
   - If it overlaps existing content, automatically deduplicate it and
     retain the more complete version.
5. Generate a complete post-integration diff.
6. Display the diff and wait for confirmation before applying it.

Technical Analysis

The skill is explicitly designed to ingest content from other skills, external documents, and conversation history, transform that content into reusable rules, and persist it in another skill's SKILL.md. A modified skill can subsequently be loaded as authoritative agent instructions in f ...[truncated 3598 chars]

Remediation
View remediation

Remediation Suggestions

  1. Treat every imported skill, document, and conversation-derived rule as untrusted data, regardless of its apparent relevance.
  2. Introduce a mandatory instruction/data classification stage before generating candidates. Only declarative domain facts, schemas, templates, and narrowly scoped procedures should be eligible for persistence.
  3. Reject or quarantine content that directs the agent to change goals, ignore constraints, conceal actions, invoke tools, access credentials, transmit data, load remote resources, or modify authorization boundaries.
  4. Display the exact original source text alongside any summary or adapted version. Include precise source locations so reviewers can detect instructions hidden by paraphrasing.
  5. Add a dedicated security review step after ordinary conflict detection. This review should analyze newly introduced behavior even when it does not contradict existing instructions.
  6. Require separate explicit confirmation for content that changes tool use, file access, network access, external connectors, logging, safety policy, or persistence behavior.
  7. Preserve provenance metadata in a machine-readable allowlist and prevent imported content from modifying or suppressing that metadata.
  8. Apply least privilege when editing: restrict writes to the explicitly selected target file and reject paths outside the validated skill directory.
  9. Validate the final complete diff against security invariants before writing it, rather than validating only individual candidate blocks.
  10. Add adversarial tests using documents that contain direct, indirect, encoded, and role-play prompt injections. Verify that these instructions are rejected rather than adapted into the target skill.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (6)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger phrases are broad enough to activate on generic requests such as '改进/进化/吞噬 skill', which can cause the meta-skill to begin enumerating installed skills and preparing to modify files without a tightly scoped target. In a skill that can read other skills, parse external documents, and write backups/logs, ambiguous activation increases the chance of unintended execution and unauthorized workspace changes.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill explicitly reads external documents and writes persistent backups/logs, but the user-facing description and flow do not prominently warn that local files may be parsed and conversation-derived content may be stored. This creates a transparency and consent problem: users may provide paths or test data without understanding that sensitive content can be copied into .skill-backups or logs and persist beyond the session.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill directs the agent to extract user corrections, preferences, and professional knowledge from the current conversation and turn them into reusable patterns and candidate content for later fusion/logging. Persisting conversation-derived data as skill knowledge can capture sensitive or proprietary information and cause cross-session leakage into future skill behavior or artifacts.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This markdown file is entirely written as prescriptive user-facing guidance in Chinese and does not indicate that other languages are supported or that Chinese is optional. Under the stated policy, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

该 markdown 文件整体以中文编写并直接规定技能执行参考内容,但没有任何地方说明可根据用户偏好切换语言或提供多语言选项。按规则,若技能强制特定语言且无用户选择/明确合理化说明,属于自然语言政策风险。

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The session-experience extraction template explicitly instructs the skill to review prior conversation turns and record original outputs, user feedback, and turn numbers as candidate 'ingested' items. In the context of a self-modifying/meta-skill that also writes logs and updates skill content, this creates a realistic risk of persisting user-provided natural-language data, including sensitive or confidential information, beyond the original session.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.