T02 · Agent Memory Poisoning
- Location
SKILL.md:78- Finding
Unverified Behavioral Inferences Written to Persistent User Memory
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 78-90
Vulnerability Type: Persistent memory poisoning through automatic profile modification
Risk Level: MediumVulnerable Code
markdown ### 2. Independent Confirmation When a pattern repeats 3+ times across different interactions, auto-confirm it: - Log to `notes/patterns.md` with `[AUTO-CONFIRMED]` tag - Update `USER.md` immediately - Mention it next conversation: "I've added X to your profile based on repeated behavior" **Auto-confirm criteria:** - Same type of content saved 3+ times (e.g., marketing frameworks) - Same intent signal repeated (e.g., always wants reminders) - Same reaction pattern (e.g., always labels overpromises as "AI porn") - Consistent preference expressed in different contexts Hidden patterns (things Enzo didn't consciously notice) are especially valuable — surface these even if auto-confirmed.Technical Analysis
The skill directs the agent to treat a behavioral inference as confirmed after three occurrences and then immediately write it to
USER.md, which functions as persistent user state. Explicit user approval is not required before this modification.Repetition alone does not establish that an observation is accurate, durable, attributable to the user, or appropriate for long-term storage. Context-specific statements, jokes, quoted material, attacker-supplied content, and interactions involving other participants could satisfy the heuristic. Once written to
USER.md, the inferred information may be treated as authoritative in future sessions.This creates a memory-poisoning boundary violation: potentially attacker-influenced conversational content is converted into trusted, cross-session profile data. The instruction to prioritize hidden patterns further increases the likelihood that information the user has never reviewed or consciously endorsed will be persisted.
Attack Path
- An attack ...[truncated 1749 chars]
- Remediation
View remediation
Remediation Suggestions
- Require explicit, informed user approval before every modification to
USER.md. - Store inferred patterns only in
notes/patterns.mduntil the user confirms them. - Label all unconfirmed entries as hypotheses rather than using
[AUTO-CONFIRMED]. - Record provenance for every observation, including timestamp, source interaction, confidence, and whether the content was directly stated by the user.
- Exclude quoted content, third-party messages, group-chat content, and other untrusted sources from confirmation heuristics.
- Treat repetition as a signal for review, not as proof of a durable preference or goal.
- Present proposed profile changes to the user with options to approve, reject, edit, or defer them.
- Implement retention and deletion controls for behavioral observations and provide a mechanism to inspect and correct stored profile data.
- Ensure downstream agents distinguish user-confirmed facts from inferred or unverified observations.
- Restrict writes to the minimum necessary fields and preserve an audit trail so unauthorized or erroneous changes can be reverted.
- Require explicit, informed user approval before every modification to
