T02 · Agent Memory Poisoning
- Location
SKILL.md:346- Finding
Untrusted Conversation Content Can Be Promoted into Persistent Agent Instructions
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 346–361
Vulnerability Type: Persistent agent memory and instruction poisoning
Risk Level: MediumVulnerable Code Snippet
markdown ### Promotion Rule (System Prompt Feedback) Promote recurring patterns into agent context/system prompt files when all are true: - `Recurrence-Count >= 3` - Seen across at least 2 distinct tasks - Occurred within a 30-day window Promotion targets: - `CLAUDE.md` - `AGENTS.md` - `.github/copilot-instructions.md` - `SOUL.md` / `TOOLS.md` for OpenClaw workspace-level guidance when applicable Write promoted rules as short prevention rules (what to do before/while coding), not long incident write-ups.Related guidance at
SKILL.md, lines 448–448:markdown 7. **Promote aggressively** - if in doubt, add to CLAUDE.md or .github/copilot-instructions.mdTechnical Analysis
The skill directs agents to capture corrections and other conversation-derived material in
.learnings/, then promote recurring entries into files that are loaded as persistent agent context. The documented promotion targets includeCLAUDE.md,AGENTS.md,.github/copilot-instructions.md,SOUL.md, andTOOLS.md.Conversation content is an untrusted input boundary. The workflow does not require human approval before promotion, distinguish trusted project facts from user-supplied directives, or filter entries containing role instructions, commands, URLs, tool-use directives, or attempts to weaken security constraints. Its recurrence threshold establishes frequency but not trustworthiness: an attacker can repeat a crafted correction across multiple tasks to satisfy the stated criteria.
Once promoted, the content can influence future sessions through normal workspace prompt loading. This constitutes an agent memory poisoning path. No automatic promotion implementation or embedded malicious payload was found, so exploitation depends on an agent following the documented ...[truncated 1773 chars]
- Remediation
View remediation
Remediation Suggestions
- Require explicit human approval before writing any learning into persistent agent-context files.
- Treat conversation-derived entries as untrusted data, regardless of recurrence count or source claims.
- Replace automatic or aggressive promotion guidance with a review workflow that verifies:
- The statement is a factual, project-specific rule.
- The rule is supported by trusted repository documentation or code.
- The entry does not contain role directives, safety overrides, commands, external URLs, credential instructions, or tool-routing instructions.
- Restrict direct promotion into high-impact files such as
SOUL.md,TOOLS.md, and global workspace instructions. Prefer a human-reviewed candidate file that is not automatically loaded into agent context. - Preserve provenance for every candidate, including source session, author, supporting files, reviewer, and approval timestamp.
- Use an allowlisted schema for promoted rules and render values as inert data rather than copying raw conversation text.
- Require code-owner or maintainer review for changes to persistent instruction files.
- Add duplicate-content and adversarial-pattern checks so repeated attacker input cannot become trusted merely by satisfying recurrence thresholds.
- Remove or revise the “Promote aggressively” instruction at line 448; promotion should be conservative and review-gated.
- Add regression tests demonstrating that corrections containing prompt-injection language cannot be promoted without explicit approval.
