T02 · Agent Memory Poisoning
- Location
SKILL.md:26- Finding
Untrusted Session Instructions Can Be Persisted as Agent Memory Without Confirmation
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md:26-40, 61-67, 76-83; related persistence criteria inreferences/learning-rules.md:5-8
Vulnerability Type:T02: Agent Memory Poisoning
Risk Level: MediumVulnerable Code Snippets
SKILL.md:26-40markdown ### Step 1: Scan the session Review the whole session for two categories: **A. Workflow preferences and corrections** - User corrections ("don't do that", "stop ...") - User-confirmed non-obvious approaches ("yes, exactly", "perfect") - Output format, communication style, and collaboration flow preferences - Tool usage guidance - `prompt-refiner` choice preferences (prefers refined vs original, wants compare-before-execute) **Judging prompt-refiner preference signals:** - Repeatedly choosing the same version → store as stable preference - One-off different choice → don't overfit, may be scenario-specific - Explicit verbal correction ("don't show me the original anymore") → promote immediately to ruleSKILL.md:61-67markdown ### Step 3: Read current CLAUDE.md Choose target file by scope: - **Global rules** → `~/.claude/CLAUDE.md` - **Project-specific rules** → project `CLAUDE.md` If a project file is needed and missing, create it.SKILL.md:76-83markdown ### Step 5: Confirm changes Show the user: - Rules to add - Rules to update (old → new) - Target file Wait for confirmation before writing, unless user says "update directly" or the skill is running from an automatic SessionEnd hook.references/learning-rules.md:5-8markdown ### High-value signals (always learn) - User explicitly corrects your behavior → immediate rule - User confirms a non-obvious approach → stable preference - Repeated pattern across 3+ interactions → stable convention - User states a project rule ("we always use X for Y") → project conventionTechnical Analysis
The Skill interpr ...[truncated 2409 chars]
- Remediation
View remediation
Remediation Suggestions
- Require explicit user confirmation before every persistent write, including writes initiated by SessionEnd hooks. Remove the automatic-confirmation bypass.
- Display the exact proposed rule, its source message, destination file, and scope before requesting approval.
- Accept learning signals only from direct user-authored messages. Exclude tool output, retrieved content, repository files, quoted text, generated prompts, and assistant paraphrases.
- Reject candidate rules that modify safety constraints, instruction precedence, permission boundaries, credential handling, confirmation requirements, or network and tool access.
- Default to project-local persistence. Require separate, explicit consent for each write to global
~/.claude/CLAUDE.md. - Record provenance metadata or maintain an auditable change log so users can identify, review, and revert learned rules.
- Validate generated rules against a restrictive allowlist of preference categories, such as formatting and non-security-sensitive coding conventions.
- Treat persisted preferences as lower-priority, untrusted guidance rather than authoritative instructions that can override system or security policies.
