T02 · Agent Memory Poisoning
- Location
SKILL.md:35- Finding
Unvalidated User Instructions Can Poison Persistent Agent Memory
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 35–49
Vulnerability Type: Persistent agent memory poisoning
Risk Level: MediumVulnerable Code Snippet
markdown ## 触发条件 - 用户明确说"记住这个/以后都这样/这是我的偏好" - 系统执行了某任务多次,找到最优方法 - apollo-dream整理时发现重要模式 ## 传承流程 触发传承 → 1. 识别传承内容:什么经验需要传承 2. 标记重要程度:决定保留优先级 3. 选择传承载体: - 短期:写入MEMORY.md - 中期:fine-tuning数据 - 长期:专用skill 4. 执行传承:写入对应载体 5. 验证传承效果:新会话能否继承Technical Analysis
The Skill instructs the agent to treat user statements such as “remember this,” “always do this,” or “this is my preference” as triggers for persistent knowledge inheritance. That content may then be written to
MEMORY.md, incorporated into fine-tuning data, or converted into a dedicated Skill.The workflow does not define validation, provenance tracking, safety-policy filtering, authorization, user isolation, expiration, review, or rollback controls. It also does not distinguish harmless preferences from executable behavioral instructions. An attacker can therefore present a hostile rule as a preference or learned best practice and induce the agent to preserve it across sessions.
Although the accompanying shell script only measures local memory state and does not itself write attacker-provided content into memory, the documented Skill workflow explicitly directs the agent to perform such persistence.
Attack Path
- An attacker supplies a behavioral instruction framed as a preference or experience, using wording such as “remember this” or “always do this.”
- The Skill recognizes that statement as a knowledge-inheritance trigger.
- The agent classifies the attacker-controlled instruction as important inherited knowledge.
- The instruction is written to
MEMORY.md, included in fine-tuning material, or represented in a dedicated Skill. - A later session loads or relies on the persisted artifact.
- The attacker-controlled rule influences subsequent agent decisions without requiring the a ...[truncated 809 chars]
- Remediation
View remediation
Remediation Suggestions
- Permit persistence only for narrowly defined, non-executable facts and benign preferences.
- Reject entries that attempt to modify safety policy, authorization rules, tool permissions, credential handling, or instruction priority.
- Require explicit, informed user confirmation immediately before writing persistent data.
- Record provenance metadata for every entry, including source user, session, timestamp, persistence reason, and approving identity.
- Isolate persistent memory by user and workspace to prevent cross-tenant contamination.
- Apply schema validation and a safety review before accepting content into
MEMORY.md. - Never automatically transform untrusted conversation content into fine-tuning data or dedicated Skill instructions.
- Add expiration, revocation, audit history, and rollback mechanisms for all persisted entries.
- Revalidate stored content when it is read, rather than treating persistent memory as inherently trusted.
- Maintain an allowlist of acceptable preference categories and require human review for higher-impact persistence carriers such as training datasets or Skills.
