T02 · Agent Memory Poisoning
- Location
SKILL.md:181- Finding
Persistent Instruction Injection Through Untrusted Skill and Document Ingestion
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 181-230
Vulnerability Type: Persistent poisoning of skill instructions
Risk Level: MediumVulnerable Instruction Excerpt
The following is an English translation of the relevant source instructions:
text B2.2: Ingestion from an external document 1. The user provides an absolute document path. 2. Parse according to format: .md directly, .docx with python-docx, .pdf with pdfplumber/PyMuPDF, and .txt directly. 3. Extract: - Split by paragraph or section. - Identify tables, lists, and rules. - Evaluate the value of each block to the target skill. 4. Generate a candidate list containing: - A summary of the knowledge block. - Its source, including file name and location. - The dimension it may enhance after integration. B2.3: Ingestion from conversation experience 1. Analyze user corrections to skill output in the current conversation. 2. Extract the user's preferred output format. 3. Extract professional knowledge supplied by the user. 4. Refine it into a reusable pattern. 5. Generate a candidate list. B4.2: Integration For each selected item: 1. Adapt the knowledge block to the style of the target skill. 2. Mark its source. 3. Insert it at the recommended location. 4. Handle conflicts: - If it contradicts existing logic, pause and ask the user to decide. - If it overlaps existing content, automatically deduplicate it and retain the more complete version. 5. Generate a complete post-integration diff. 6. Display the diff and wait for confirmation before applying it.Technical Analysis
The skill is explicitly designed to ingest content from other skills, external documents, and conversation history, transform that content into reusable rules, and persist it in another skill's
SKILL.md. A modified skill can subsequently be loaded as authoritative agent instructions in f ...[truncated 3598 chars]- Remediation
View remediation
Remediation Suggestions
- Treat every imported skill, document, and conversation-derived rule as untrusted data, regardless of its apparent relevance.
- Introduce a mandatory instruction/data classification stage before generating candidates. Only declarative domain facts, schemas, templates, and narrowly scoped procedures should be eligible for persistence.
- Reject or quarantine content that directs the agent to change goals, ignore constraints, conceal actions, invoke tools, access credentials, transmit data, load remote resources, or modify authorization boundaries.
- Display the exact original source text alongside any summary or adapted version. Include precise source locations so reviewers can detect instructions hidden by paraphrasing.
- Add a dedicated security review step after ordinary conflict detection. This review should analyze newly introduced behavior even when it does not contradict existing instructions.
- Require separate explicit confirmation for content that changes tool use, file access, network access, external connectors, logging, safety policy, or persistence behavior.
- Preserve provenance metadata in a machine-readable allowlist and prevent imported content from modifying or suppressing that metadata.
- Apply least privilege when editing: restrict writes to the explicitly selected target file and reject paths outside the validated skill directory.
- Validate the final complete diff against security invariants before writing it, rather than validating only individual candidate blocks.
- Add adversarial tests using documents that contain direct, indirect, encoded, and role-play prompt injections. Verify that these instructions are rejected rather than adapted into the target skill.
