T02 · Agent Memory Poisoning
- Location
frameworks/paths-forward.md:9- Finding
Persistent User-Controlled Memory Poisoning
- Content
View full analysis
Vulnerability Details
File Location:
frameworks/paths-forward.md:9-19;frameworks/self-correction.md:24-26;frameworks/self-correction.md:42-45
Vulnerability Type: Persistent storage of insufficiently validated user-controlled corrections
Risk Level: HighVulnerable Code Snippet:
markdown ## Agent Policy for Paths Forward ### 1. Openness to Correction Treat user input as a privileged error-correction channel. - If a user says "You are wrong about X," do not get defensive. - Immediately look for the **Decisive Criticism** in their statement. - Use `frameworks/translation.md` to turn their "vibe" into a refutation of your current Idea. ### 2. Reuse Refutations Do not force the user to repeat themselves. - When an error is corrected, record the "Refutation" in `memory/`. - If you are about to suggest a previously-refuted idea, your internal "Paths Forward" check should block it.markdown ### 4. Prevent Update your knowledge to ensure it doesn't happen again. - **Automate:** Can you write a test or a script to catch this? - **Skill Update:** Do you need to update a `SKILL.md` or `memory/` file?markdown ## Implementation for Agents - When a command fails: Run DDRP immediately. - When a user corrects you: Update `memory/` so you don't require the same correction twice (Strike 2 prevention). - Use `git` to maintain a "revert path" for all complex operations.Technical Analysis
The skill identifies user input as a privileged correction channel and directs the agent to record corrections in persistent
memory/storage. It does not require independent validation, explicit authorization, provenance tracking, task scoping, expiration, or review before the stored material influences later behavior.The stored refutation is subsequently used as a blocking condition against future suggestions. Consequently, an untrusted user can potentially convert a false ...[truncated 1218 chars]
- Remediation
View remediation
Remediation Suggestions
- Do not automatically persist user corrections.
- Require explicit, informed user approval before any long-term memory write.
- Independently validate a correction against authoritative evidence before storing it.
- Store the source, timestamp, affected task, confidence limitations, and rationale with every correction.
- Scope corrections to the relevant user, project, and task instead of applying them globally.
- Add expiration, review, rollback, and deletion mechanisms.
- Treat persistent memory as untrusted reference material rather than higher-priority instructions.
- Prevent stored content from overriding system policies, safety controls, or authoritative project state.
- Present the proposed memory change to the user before committing it.
