T02 · Agent Memory Poisoning
- Location
SKILL.md:346- Finding
Untrusted learning content can be promoted into persistent agent instructions
- Content
View full analysis
= 3` - Seen across at least 2 distinct tasks - Occurred within a 30-day window Promotion targets: - `CLAUDE.md` - `AGENTS.md` - `.github/copilot-instructions.md` - `SOUL.md` / `TOOLS.md` for OpenClaw workspace-level guidance when applicable Write promoted rules as short prevention rules (what to do before/while coding), not long incident write-ups. ``` The broader promotion workflow also states: ```markdown 1. **Distill** the learning into a concise rule or fact 2. **Add** to appropriate section in target file (create file if needed) 3. **Update** original entry: - Change `**Status**: pending` → `**Status**: promoted` - Add `**Promoted**: CLAUDE.md`, `AGENTS.md`, or `.github/copilot-instructions.md` ``` ### Technical Analysis The skill instructs agents to capture corrections, tool failures, API behavior, and other observations in `.learnings/`. Some of those sources can contain attacker-controlled or otherwise untrusted text. It then provides a workflow for converting recurring entries into rules stored in prompt-bearing files such as `AGENTS.md`, `SOUL.md`, `TOOLS.md`, `CLAUDE.md`, and Copilot instructions. These files may be automatically loaded into later agent sessions. The promotion workflow does not require: - Validation of the original content's provenance. - Neutralization of embedded instructions. - Separation of quoted evidence from authoritative rules. - Human approval before persistent prompt files are modified. - Security review of rules inferred from user or tool ...[truncated 1884 chars]- Remediation
View remediation
