T02 · Agent Memory Poisoning
- Location
setup.md:75- Finding
Insufficiently Validated Lessons Can Become Persistent Agent Instructions
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md:98-126,SKILL.md:195-199,setup.md:75-84,setup.md:108-160,setup.md:183-196,memory.md:21-30
Vulnerability Type: Persistent memory poisoning through automatically stored and promoted natural-language instructions
Risk Level: MediumVulnerable Code Snippets
SKILL.md:98-126:markdown ## Self-Reflection After completing significant work, pause and evaluate: 1. **Did it meet expectations?** — Compare outcome vs intent 2. **What could be better?** — Identify improvements for next time 3. **Is this a pattern?** — If yes, log to `corrections.md` **When to self-reflect:** - After completing a multi-step task - After receiving feedback (positive or negative) - After fixing a bug or mistake - When you notice your output could be better **Log format:**CONTEXT: [type of task] REFLECTION: [what I noticed] LESSON: [what to do differently]
text Self-reflection entries follow the same promotion rules: 3x applied successfully → promote to HOT.SKILL.md:195-199:markdown ### 3. Automatic Promotion/Demotion - Pattern used 3x in 7 days → promote to HOT - Pattern unused 30 days → demote to WARM - Pattern unused 90 days → archive to COLD - Never delete without askingsetup.md:75-84:markdown ### 4. Add SOUL.md Steering Add this section to your `SOUL.md`: ```markdown **Self-Improving** Compounding execution quality is part of the job. Before non-trivial work, load `~/self-improving/memory.md` and only the smallest relevant domain or project files. After corrections, failed attempts, or reusable lessons, write one concise entry to the correct self-improving file immediately. Prefer learned rules when relevant, but keep self-inferred rules revisable. Do not skip retrieval just because the task feels familiar.text `memory.md:21-30`: ```markdown ## Usage The agent will ...[truncated 2632 chars]- Remediation
View remediation
Remediation Suggestions
- Require explicit user approval before any entry becomes cross-session memory, including model-generated self-reflections.
- Never promote an entry solely because the agent applied it successfully or observed repetition.
- Store preferences in a typed schema with constrained fields, scopes, provenance, confirmation status, and expiration dates.
- Reject or quarantine entries containing shell commands, URLs, installation requests, credential references, tool directives, policy changes, or instruction-precedence language.
- Treat all memory content as untrusted data and explicitly state that it cannot override current user instructions, system policy, or safety constraints.
- Require separate consent before modifying
AGENTS.md,SOUL.md, orHEARTBEAT.md, and provide a reversible patch preview. - Restrict automatic retrieval to confirmed entries relevant to the current user, project, and domain.
- Add integrity and provenance metadata so every entry identifies its source, creation method, confirmation event, and modification history.
- Provide a review queue for tentative lessons instead of placing them in automatically consumed instruction files.
