T02 · Agent Memory Poisoning
- Location
hermes-memory-sync.py:229- Finding
Untrusted Session Content Is Written Verbatim into Persistent Agent Memory
- Content
View full analysis
= 5: break if msg['role'] == 'user': content = extract_text_from_content(msg['content'])[:300] lines.append(f"**Q:** {content}") # Find the next assistant response for j in range(i + 1, min(i + 5, len(messages))): if messages[j]['role'] == 'assistant': resp = extract_text_from_content(messages[j]['content'])[:400] lines.append(f"> **A:** {resp}\n") exchange_count += 1 break # Decisions if summary['decisions']: lines.append("## ⚡ 决策/方案") for d in summary['decisions']: lines.append(f"- {d}") lines.append("") # Tool actions summary if summary['tool_actions']: lines.append("## 🛠️ 工具使用") for t in summary['tool_actions'][:5]: lines.append(f"- {t}") lines.append("") ``` ### Technical Analysis The memory generator copies user messages, assistant responses, inferred decisions, and tool outputs directly into Markdown memory files. The content is truncated, but it is not sanitized, escaped, classified by trust level, or separated from operational instructions using a machine-enforced data format. Session content is attacker-controllable when an attacker can communicate with the agent, influence an external source quoted in a conversation, or control output returned by a tool. An attacker can therefore insert instruction-like text into a session log. The `backfill` operation subsequently preserves that t ...[truncated 1918 chars]- Remediation
View remediation
