T02 · Agent Memory Poisoning
- Location
SKILL.md:21- Finding
Untrusted Transcript Instructions Can Be Promoted into Persistent Agent Memory
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md:21-27andreferences/principles.md:4-10
Vulnerability Type: T02: Agent Memory Poisoning
Risk Level: HighComplete Code Snippet (
SKILL.md:21-27):markdown 2. **Principles** (see references/principles.md): - Delete casual/heartbeats/repeated tools. - Keep: disciplines, configs, skills learned/install, explicit "remember", decisions. 3. **Process**: - Read raw. - Extract key. - Write summary to memory/YYYY-MM-DD-summary.md or MEMORY.md.Complete Code Snippet (
references/principles.md:4-10):markdown - **Keep**: - Explicit instructions ("do X"). - Disciplines/rules (e.g. skill-vetter first). - Configs (API keys, models). - Skills learned (install/how). - User "remember this" info. - Decisions/dates/people/preferences/todos.Technical Analysis
The workflow treats chat transcripts and session histories as source material, explicitly preserves instructions, disciplines, rules, and “remember this” statements, and then writes the extracted content to persistent memory files.
Transcripts can contain attacker-controlled or otherwise untrusted text. The documented process does not require provenance checks, trust classification, authenticated-user confirmation, or separation between quoted historical content and authoritative behavioral instructions. Consequently, malicious text embedded in a transcript can be transformed from untrusted conversation data into a durable rule that future agent sessions may treat as trusted memory.
Attack Path
- An attacker places a plausible instruction, discipline, or “remember this” statement in a transcript available to the agent.
- A user invokes the chat-refiner skill against that transcript or its associated session history.
- The skill follows its documented requirement to retain explicit instructions and rules.
- The malicious instruction is written to
MEMORY.mdor `memory/YYYY-MM-DD-summary.md ...[truncated 625 chars]
- Remediation
View remediation
Remediation Suggestions
- Do not automatically promote transcript instructions, disciplines, or rules into authoritative memory.
- Treat all transcript-derived instructions as untrusted quoted data and preserve their source, author, session identifier, and timestamp.
- Require explicit confirmation from the authenticated user before persisting any behavioral instruction.
- Reject or quarantine content that attempts to alter safety constraints, tool permissions, trust boundaries, or memory-processing rules.
- Separate factual user preferences from executable or imperative instructions using a structured schema.
- Mark generated summaries as non-authoritative context so future agents cannot interpret them as higher-priority instructions.
- Add a review stage that displays proposed memory additions before writing them.
