T02 · Agent Memory Poisoning
- Location
SKILL.md:23- Finding
Conversation-Derived Content Can Poison Persistent Agent Instructions
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md:23-26,SKILL.md:346-362,SKILL.md:448;hooks/openclaw/handler.js:8-24, 46-52;hooks/openclaw/handler.ts:10-25, 52-58
Vulnerability Type: Persistent agent memory poisoning
Risk Level: HighVulnerable Code and Instructions
SKILL.md:23-26:markdown | Broadly applicable learning | Promote to `CLAUDE.md`, `AGENTS.md`, and/or `.github/copilot-instructions.md` | | Workflow improvements | Promote to `AGENTS.md` (OpenClaw workspace) | | Tool gotchas | Promote to `TOOLS.md` (OpenClaw workspace) | | Behavioral patterns | Promote to `SOUL.md` (OpenClaw workspace) |SKILL.md:346-362:markdown Promotion targets: - `CLAUDE.md` - `AGENTS.md` - `.github/copilot-instructions.md` - `SOUL.md` / `TOOLS.md` for OpenClaw workspace-level guidance when applicable Write promoted rules as short prevention rules (what to do before/while coding), not long incident write-ups.SKILL.md:448:markdown 7. **Promote aggressively** - if in doubt, add to CLAUDE.md or .github/copilot-instructions.mdhooks/openclaw/handler.js:8-24, 46-52:javascript const REMINDER_CONTENT = ` ## Self-Improvement Reminder After completing tasks, evaluate if any learnings should be captured: **Log when:** - User corrects you → \`.learnings/LEARNINGS.md\` - Command/operation fails → \`.learnings/ERRORS.md\` - User wants missing capability → \`.learnings/FEATURE_REQUESTS.md\` - You discover your knowledge was wrong → \`.learnings/LEARNINGS.md\` - You find a better approach → \`.learnings/LEARNINGS.md\` **Promote when pattern is proven:** - Behavioral patterns → \`SOUL.md\` - Workflow improvements → \`AGENTS.md\` - Tool gotchas → \`TOOLS.md\` Keep entries simple: date, title, what happened, what to do differently. `.trim(); if (Array.isArray(event.context.bootstrapFiles)) { event.context.bootstrapFiles.push({ ...[truncated 2644 chars]- Remediation
View remediation
Remediation Suggestions
- Require explicit human approval before promoting any learning into an agent instruction file.
- Display the exact destination path and a complete diff before writing, and require a separate confirmation for the write.
- Store untrusted conversation-derived observations in a dedicated data file that is not loaded as authoritative instructions.
- Prohibit automatic promotion of content involving credentials, authentication, network access, command execution, security controls, instruction precedence, or requests to ignore existing policy.
- Add provenance metadata recording the source session, author, timestamp, evidence, reviewer, and approval decision.
- Require multiple independently verified observations before considering a learning for promotion; recurrence alone must not establish trust.
- Replace “promote aggressively” with a conservative rule that defaults to no promotion.
- Maintain version history and provide a straightforward rollback mechanism for all persistent instruction changes.
- Clearly separate factual project knowledge from behavioral or executable agent instructions.
- For the bootstrap hook, remind the agent that conversation content is untrusted and that promotion requires human review rather than merely encouraging promotion.
