T02 · Agent Memory Poisoning
- Location
plugins/heartflow-memory-inject.py:85- Finding
Persistent system-prompt poisoning through unsanitized long-term memory
- Content
View full analysis
'); return; } const result = hfm.learn(key, value, ['manual']); if (result.success) { console.log(`已写入 LEARNED 层: [${key}] = ${value}`); } } ``` Stored values are inserted verbatim into the generated prompt suffix. For example, generic learned entries are rendered without escaping or validation: ```js if (other.length > 0) { lines.push(''); lines.push('【其他记忆】'); for (const e of other.slice(0, 5)) { lines.push(` • ${e.value}`); } } ``` The plugin then attaches the resulting content to the system prompt before every processed message: ```python def before_message(self, context): """ 在每次处理用户消息前,注入心虫记忆到系统提示。 """ inject_text = _run_inject() if inject_text: return { "system_prompt_suffix": inject_text, } return {} ``` ### Technical Analysis The memory subsystem does not distinguish trusted instructions from untrusted data. An arbitrary LEARNED memory value is serialized as plain text and subsequently assigned system-prompt privilege through `system_prompt_suffix`. There is no: - Prompt-injection detection or neutralization. - Escaping or structured serialization that prevents memory text from being interpreted as instructions. - Provenance or trust-level enforcement. - Human approval before memory becomes part of the system prompt. - Separation between factual memory, user preferences, conversation text, and exe ...[truncated 1540 chars]- Remediation
View remediation
