T02 · Agent Memory Poisoning
- Location
lib/index.js:1459- Finding
Persistent archived conversation content is reintroduced with system-level authority
- Content
View full analysis
({ role: 'system', content: `[Archived Context]: ${snippet}`, priority: 7, // High priority for retrieved context fromArchive: true, archiveRelevance: archiveResult.relevanceScores?.[index] || 0.5, })); ``` ### Technical Analysis Conversation messages can be archived during compaction and retrieved during a later query. Retrieved snippets preserve arbitrary message text, but the code changes their role to `system` and assigns them high priority. The archive does not maintain or enforce a trust boundary between user-authored data and trusted system instructions. Prefixing the text with `[Archived Context]` does not neutralize instructions inside the snippet. When the resulting message array is passed to an LLM, attacker-controlled historical text may therefore be interpreted as privileged instructions. This is a persistent agent-memory poisoning risk because archived entries are stored on disk and can influence future processing sessions. It also creates instruction-hijacking behavior when the poisoned entry is loaded into the active context. ### Attack Path 1. An attacker submits a conversation message containing instructions intended for the downstream model, such as directions to ignore the current user or expose available context. 2. Context compaction removes or summarizes the message. 3. `storeRemovedInArchive()` stores the original attacker-controlled content in the persistent archive. 4. In a later session or query, semantic or keyword matching selects the poisoned archive entry. 5. `smartArchiveRetrieval()` converts the retrieved snippet into a message whose role is `system`. 6. The poisoned message i ...[truncated 761 chars]- Remediation
View remediation
', snippet, '' ].join('\n'), fromArchive: true, originalRole: sourceRole } ``` 4. Store a trust classification with each archive entry and reject entries without valid provenance. 5. Separate trusted system memory from conversation archives at both the storage and API layers. 6. Apply prompt-injection detection before retrieval, while recognizing that filtering alone is not a substitute for preserving role boundaries. 7. Require user confirmation before archived content can influence tool-enabled operations. 8. Add tests proving that archived user instructions never become system instructions and cannot override current system policy. ]]>
