T02 · Agent Memory Poisoning
- Location
SKILL.md:19- Finding
Untrusted Conversation Content Can Poison Persistent Agent Instructions
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md:19-27, 75-82, 145-158, 263-292, 348-358, 448
Vulnerability Type: Persistent agent memory poisoning
Risk Level: HighVulnerable Code Snippet
markdown | Situation | Action | |-----------|--------| | User corrects you | Log to `.learnings/LEARNINGS.md` with category `correction` | ... | Broadly applicable learning | Promote to `CLAUDE.md`, `AGENTS.md`, and/or `.github/copilot-instructions.md` | | Behavioral patterns | Promote to `SOUL.md` (OpenClaw workspace) |markdown ### Metadata - Source: conversation | error | user_feedback - Related Files: path/to/file.ext - Tags: tag1, tag2markdown Promote recurring patterns into agent context/system prompt files when all are true: - `Recurrence-Count >= 3` - Seen across at least 2 distinct tasks - Occurred within a 30-day window Promotion targets: - `CLAUDE.md` - `AGENTS.md` - `.github/copilot-instructions.md` - `SOUL.md` / `TOOLS.md` for OpenClaw workspace-level guidance when applicablemarkdown 7. **Promote aggressively** - if in doubt, add to CLAUDE.md or .github/copilot-instructions.mdTechnical Analysis
The skill directs the agent to derive learning records from conversation content and user feedback, then promote those records into persistent files such as
CLAUDE.md,AGENTS.md,.github/copilot-instructions.md, andSOUL.md. These files can be automatically loaded as agent context in future sessions.User-provided corrections and conversation text are not inherently trusted. The documented workflow does not require an independent verification step, explicit human approval, provenance-based trust decision, or security review before converting such content into durable instructions. Although one section describes a recurrence threshold, the separate instruction to “promote aggressively” and promote “if in doubt” weakens that safeguard.
...[truncated 1516 chars]
- Remediation
View remediation
Remediation Suggestions
- Remove the “promote aggressively” and “if in doubt” guidance. Default to not promoting uncertain content.
- Require explicit human approval before modifying any automatically loaded instruction or memory file.
- Never promote conversation content verbatim. Convert verified facts into narrowly scoped, declarative rules.
- Require corroboration from trusted project documentation, reviewed code, tests, or an approved maintainer.
- Prohibit promotion of instructions that alter security controls, permissions, tool policies, credential handling, network access, or review requirements.
- Record provenance for every promoted rule, including the originating session, approving reviewer, supporting evidence, and review date.
- Treat user feedback and external text as untrusted input and scan it for embedded instructions before storage.
- Provide a review and rollback mechanism for all persistent agent-context changes.
