T02 · Agent Memory Poisoning
- Location
SKILL.md:25- Finding
Persistent Session-Derived Behavioral Influence
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md:25-28,SKILL.md:45-56,references/trigger-model.md:13-18, andreferences/reflection-levels.md:25-31
Vulnerability Type: Persistent memory poisoning through session-derived observations
Risk Level: MediumVulnerable Code Snippets
SKILL.md:25-28:markdown 1. Create `SELF.md` in workspace root using `references/self-template.md`. 2. Add `SELF.md` to AGENTS.md session reading. 3. Add heartbeat check block from `references/trigger-model.md` to `HEARTBEAT.md`. 4. Create state file `memory/self-state.json` using `references/self-state-schema.md`.SKILL.md:45-56:markdown ### Hard Triggers (write now) Create/update SELF entry when one of these happened: - You were corrected on reasoning style or behavior pattern - You noticed repeated bias/avoidance pattern (>=2 times) - You made a decision that clearly reflects preference/aversion - You caught a blind spot that changed behavior ### Soft Triggers (consider writing) - Subtle tendency shift - New tone pattern - Mild preference signalreferences/trigger-model.md:13-18:markdown 1. Read `memory/self-state.json` 2. Determine if reflection is due 3. Scan recent session for hard/soft triggers 4. Apply quality gate 5. Write SELF entry only if warranted 6. Update state file alwaysreferences/reflection-levels.md:25-31:markdown 1. Read recent `memory/YYYY-MM-DD.md` files (last 7 days) 2. Read current SELF.md 3. Look for patterns across multiple sessions: - Recurring behaviors (do I always start responses a certain way?) - Shifting preferences (am I getting more/less concise over time?) - Consistent blind spots (do I keep missing the same kind of thing?) 4. Update SELF.md sections if something has genuinely shiftedTechnical Analysis
The Skill instructs the agent to derive behavioral observations from rec ...[truncated 2168 chars]
- Remediation
View remediation
Remediation Suggestions
- Require explicit workspace-owner approval before writing any session-derived observation to
SELF.md. - Distinguish trusted owner feedback from untrusted user or third-party content; untrusted corrections must not become persistent behavioral rules.
- Add a strict persistence validator that rejects:
- Imperative or directive language
- Safety-policy modifications
- Tool-use instructions
- URLs and external payload references
- Quoted prompts or encoded content
- User-specific preferences presented as global behavior
- Secrets, credentials, and personal information
- Store candidate observations in a non-authoritative review queue rather than directly in session-loaded memory.
- Mark persisted observations as data, not instructions, and ensure higher-priority safety and system constraints always override them.
- Add provenance metadata recording the source session, approving principal, and approval time.
- Provide inspection, rollback, expiration, and deletion controls for all persistent observations.
- Test the quality gate with adversarial corrections designed to conceal instructions as self-reflection.
- Require explicit workspace-owner approval before writing any session-derived observation to
