T02 · Agent Memory Poisoning
- Location
SKILL.md:27- Finding
Automated Promotion of Untrusted Session Content into Persistent Agent Memory
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md:27-32andSKILL.md:40-57
Vulnerability Type: Agent memory poisoning through automatic memory consolidation
Risk Level: MediumRelevant Snippets:
markdown Dreaming automatically: 1. **Light phase** — stages recent daily notes and session transcripts 2. **REM phase** — extracts patterns and recurring themes 3. **Deep phase** — promotes strong signals into `MEMORY.md` It runs as a background cron job tied to heartbeat.json5 { plugins: { entries: { "memory-core": { config: { dreaming: { enabled: true, // Optional: custom schedule (default: 03:00 daily) frequency: "0 3 * * *", timezone: "Asia/Jerusalem", }, }, }, }, }, }Technical Analysis
The documented configuration creates an automated path from recent notes and session transcripts into the persistent
MEMORY.mdfile. Promotion is based on extracted patterns and recurring signals, but the Skill does not document any trust-boundary validation, provenance enforcement, prompt-injection filtering, content-type restrictions, or mandatory human approval before the Deep phase writes persistent memory.An attacker who can influence session transcripts or daily notes could repeatedly introduce false facts, behavioral directives, or instruction-like content. Recurrence-based consolidation may then interpret that content as a strong signal and promote it into long-term memory. Once stored, the poisoned content can continue to affect later sessions that use
MEMORY.md.This issue does not grant operating-system privileges or independently provide arbitrary code execution. Its scope is the persistent memory and subsequent behavior of the affected OpenClaw agent.
Attack Path
- An administrator enables Dreaming using the documented configuration.
...[truncated 1150 chars]
- Remediation
View remediation
Remediation Suggestions
- Require explicit human approval for every Deep-phase change to
MEMORY.md, particularly when source material includes session transcripts. - Present the exact proposed diff, source excerpts, source timestamps, and originating sessions before promotion.
- Treat transcripts and user-authored notes as untrusted input. Detect and reject instruction-like content, prompt-injection patterns, executable commands, and attempts to modify agent policy.
- Preserve provenance for every promoted item and distinguish verified user preferences from unverified conversational claims.
- Use an allowlist of data classes that may enter persistent memory, such as explicitly confirmed preferences or administrator-approved facts.
- Do not use repetition alone as evidence of trustworthiness. Apply source diversity, confidence, and explicit verification requirements.
- Default automatic promotion to preview-only mode and require a separate approval action before persistent writes.
- Maintain versioned backups and an auditable promotion log so poisoned entries can be identified and rolled back reliably.
- Add per-source exclusion controls so sensitive, external, anonymous, or otherwise untrusted conversations cannot participate in memory consolidation.
- Require explicit human approval for every Deep-phase change to
