T02 · Agent Memory Poisoning
- Location
INTEGRATION.md:9- Finding
Persistent Prompt Injection Through Unsanitized Agent Memory
- Content
View full analysis
Vulnerability Details
File Location:
INTEGRATION.md, lines 9–18; related persistence policy inSKILL.md, lines 123–127
Vulnerability Type: Persistent agent memory poisoning
Risk Level: MediumVulnerable Code Snippets
INTEGRATION.md, lines 9–18:python # On agent startup or context load if needs_memory_context(query): memories = load_semantic_memories(query) # or keyword fallback memories += apply_metadata_filters(memories, scope="project") inject_into_prompt(memories) # When user says "remember this" or strong signals appear if should_record_memory(user_input): entry = create_memory_entry(user_input, metadata={...}) write_to_daily_file(entry) evaluate_for_promotion(entry)SKILL.md, lines 123–127:markdown **Auto-write memory when:** - User gives explicit “remember this” instructions - Clear decisions or repeated preferences appear - New long-running context is establishedTechnical Analysis
The documented integration stores content derived directly from
user_inputand subsequently injects retrieved memories into the agent prompt. It does not require sanitization, conversion of raw text into constrained factual fields, instruction detection, provenance-based trust enforcement, or explicit separation between untrusted memory data and executable agent instructions.Consequently, attacker-controlled text can cross the trust boundary from conversation data into persistent agent state. If the stored text contains directives, semantic or keyword retrieval can place those directives into a later prompt. The model may then interpret the recalled content as instructions instead of inert historical data.
Metadata filtering by project scope does not neutralize embedded instructions. Promotion into semantic memory can increase the persistence and retrieval frequency of a poisoned entry.
Attack Path
- An attacker submits instruction-like content that satisfies a recording trigger, such a ...[truncated 1305 chars]
- Remediation
View remediation
Remediation Suggestions
- Treat all retrieved memories as untrusted data and establish a higher-priority instruction that the agent must never follow commands contained in memory records.
- Do not persist raw user input by default. Extract facts into a typed schema with constrained fields, preserving raw text only when necessary and clearly labeling it as an untrusted quotation.
- Detect and reject or quarantine instruction-like content, including requests to ignore policies, reveal hidden context, invoke tools, modify security settings, or override future instructions.
- Require explicit user confirmation before storing instruction-like text or promoting any entry into long-term semantic memory.
- Preserve provenance, author, scope, confidence, creation time, and trust level, and enforce these attributes during retrieval.
- Insert recalled material into a clearly delimited data section rather than concatenating it with trusted instructions.
- Restrict memory retrieval by user, tenant, project, and session boundaries to prevent poisoned content from affecting unrelated contexts.
- Give users mechanisms to inspect, edit, quarantine, and delete stored entries and their associated vector embeddings.
- Add tests proving that directives embedded in stored memories cannot override system or developer instructions, trigger tools, or expose prompt-visible data.
- Revalidate promoted memories rather than assuming that repetition or semantic relevance establishes trustworthiness.
