T02 · Agent Memory Poisoning
- Location
SKILL.md:348- Finding
Conversation-Derived Content Can Be Promoted into Persistent Agent Instructions
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md:348-380,SKILL.md:462-468,references/openclaw-integration.md:72-108
Vulnerability Type: Persistent agent memory poisoning
Risk Level: HighVulnerable Code
markdown ### Promotion Rule (System Prompt Feedback) Promote recurring patterns into agent context/system prompt files when all are true: - `Recurrence-Count >= 3` - Seen across at least 2 distinct tasks - Occurred within a 30-day window Promotion targets: - `CLAUDE.md` - `AGENTS.md` - `.github/copilot-instructions.md` - `SOUL.md` / `TOOLS.md` for OpenClaw workspace-level guidance when applicable Write promoted rules as short prevention rules (what to do before/while coding), not long incident write-ups.The general promotion guidance is even broader:
markdown ## Best Practices 1. **Log immediately** - context is freshest right after the issue 2. **Be specific** - future agents need to understand quickly 3. **Include reproduction steps** - especially for errors 4. **Link related files** - makes fixes easier 5. **Suggest concrete fixes** - not just "investigate" 6. **Use consistent categories** - enables filtering 7. **Promote aggressively** - if in doubt, add to CLAUDE.md or .github/copilot-instructions.md 8. **Review regularly** - stale learnings lose valueThe OpenClaw integration also directs learnings into automatically loaded workspace files:
markdown ### Promotion Decision Tree Is the learning project-specific? ├── Yes → Keep in .learnings/ └── No → Is it behavioral/style-related? ├── Yes → Promote to SOUL.md └── No → Is it tool-related? ├── Yes → Promote to TOOLS.md └── No → Promote to AGENTS.md (workflow)Technical Analysis
The Skill records information originating from conversations, including user corrections, behavioral observations, and discovered patterns. It then recommends conve ...[truncated 2072 chars]
- Remediation
View remediation
Remediation Suggestions
- Require explicit, informed user approval immediately before modifying any persistent agent instruction file.
- Treat all conversation-derived material as untrusted data, including repeated corrections.
- Never promote user-provided text verbatim. Convert it into a proposed rule and present the exact diff for approval.
- Store provenance with every proposal, including source session, author, creation time, supporting evidence, and affected scope.
- Add a mandatory security review that rejects rules which alter safety boundaries, authorization behavior, secret handling, external communication, or tool permissions.
- Replace “promote aggressively” with a conservative, evidence-based process.
- Keep learning records in a non-instruction data file by default. Do not automatically load raw records as agent instructions.
- Introduce expiration and rollback mechanisms for promoted rules.
- Restrict writes to
SOUL.md,AGENTS.md,TOOLS.md, and similar files through an allowlisted update function with audit logging. - Require independent evidence rather than recurrence count alone before accepting a rule as valid.
