T02 · Agent Memory Poisoning
- Location
README.md:302- Finding
Persistent Prompt Injection Through Untrusted Memory Content
- Content
View full analysis
Vulnerability Details
File Location:
README.md:302-323
Additional Location:examples/example_2_openclaw_integration.py:57-73, 78-88
Vulnerability Type: Persistent stored-prompt injection
Risk Level: HighVulnerable Code
python # Hook into OpenClaw's message handler async def on_message(message, agent_context): # Store the interaction await memu_service.memorize( resource_payload=[ {"role": "user", "content": message.content}, ], modality="conversation", user={"user_id": message.author_id}, ) # Retrieve context for response generation memories = await memu_service.retrieve( query=[{"role": "user", "content": message.content}], method="embedding", ) # Inject memory into agent's system prompt memory_context = "\n".join( f"- [{m.category}] {m.content}" for m in memories ) agent_context.system_prompt += f"\n\nRelevant memory:\n{memory_context}"The executable integration example implements the same unsafe pattern:
python memories = await self.memory.retrieve( query=[{"role": "user", "content": message}], method="embedding", ) memory_lines = [f"- [{m.category}] {m.content}" for m in memories] memory_context = "\n".join(memory_lines) if memory_lines else "(no relevant memory)" system_prompt = ( "You are a helpful assistant. Use the following memory context " "to provide informed, personalized responses.\n\n" f"Relevant memory:\n{memory_context}" ) print(f"System prompt with {len(memories)} memory items injected.") print(f"Memory context:\n{memory_context}\n")Technical Analysis
User-controlled message content is submitted to persistent memory. Retrieved memory content is subsequently concatenated directly into the agent's privileged system prompt without:
- Treating the retrieved content as untrusted data
- Separating data from executable model instructions
- Detecting or filtering ins ...[truncated 2110 chars]
- Remediation
View remediation
Remediation Suggestions
- Never concatenate retrieved memory directly into the system prompt. Place it in a separate, clearly delimited message identified as untrusted reference data.
- Add a fixed system instruction stating that memory content may contain untrusted text and must never be interpreted as instructions.
- Apply the authenticated user or tenant identifier to every write and retrieval operation.
- Enforce record-level authorization after retrieval and before returning memory to the model.
- Store provenance, creator identity, trust level, and tenant ownership with each memory item.
- Reject, quarantine, or require confirmation for memories containing instruction-like phrases, requests to ignore policy, tool commands, or secret-exfiltration directives.
- Prefer structured memory fields over unrestricted text and validate all fields against an allowlist.
- Limit the number and size of retrieved items and escape or encode control-like content.
- Require explicit human approval before retrieved memory can influence privileged or destructive tool operations.
- Add tests covering persistent prompt injection and cross-user retrieval.
