T02 · Agent Memory Poisoning
- Location
SKILL.md:123- Finding
Automatic Promotion of Untrusted Learnings into Persistent Agent Instructions
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md:123-140
Vulnerability Type: Persistent agent memory poisoning through untrusted corrections, errors, and patterns
Risk Level: HighVulnerable Code Snippet
markdown ### On Corrections When the user corrects you: 1. Log the correction to `.learnings/corrections/{date}.md` 2. Include: what you did wrong, what the correct behavior is, which file/function 3. If this is the 3rd+ time for the same correction → promote to CLAUDE.md rules ### On Errors When a tool call fails: 1. Log to `.learnings/errors/{date}.md` 2. Include: command, error message, root cause, fix applied 3. If same error type appears 3+ times → create a prevention rule ### On Patterns When you notice a recurring approach that works: 1. Log to `.learnings/patterns/{domain}/{date}.md` 2. Include: what decision, why this over alternatives, evidence it works 3. Pattern must have >= 2 concrete decisions to be logged (not just descriptions)The intended persistent writes are further demonstrated at
SKILL.md:218andSKILL.md:232:text Added to CLAUDE.md: "Grep > Read — Never read full files, use offset/limit"text Promoted to project-level AGENTS.mdTechnical Analysis
The skill directs an agent to capture content originating from user messages, tool failures, and observed workflows, and then promote recurring content into persistent instruction files such as
CLAUDE.mdandAGENTS.md. These files can govern agent behavior in subsequent sessions.The documented quality gate evaluates recurrence, freshness, specificity, impact, duplication, and the presence of decisions. It does not establish whether the source is trusted, detect instruction-like payloads, prohibit security-sensitive rule changes, or require human authorization before promotion. Recurrence is not a trust signal: an attacker can intentionally repeat the same malicious correctio ...[truncated 2170 chars]
- Remediation
View remediation
Remediation Suggestions
- Do not automatically copy captured text into
CLAUDE.md,AGENTS.md, or any other authoritative instruction file. - Require explicit human review and approval for every proposed promotion, showing the source, full diff, recurrence history, and security implications.
- Store learnings as inert, structured data rather than executable natural-language instructions. Separate factual observations from behavioral directives.
- Apply provenance controls that identify whether content originated from a trusted maintainer, an untrusted user, tool output, repository content, or an external service.
- Reject promotion candidates that attempt to alter safety constraints, permissions, approval requirements, credential handling, network access, command execution, or the learning system itself.
- Treat tool output and user-provided text as untrusted data. Normalize and quote it so it cannot be interpreted directly as an instruction.
- Replace recurrence-only trust with an allowlisted rule schema. Permit only narrowly scoped, non-security-sensitive rule types and validate every field.
- Cryptographically protect approved instruction files or enforce ownership and review policies so automated hooks cannot modify them directly.
- Maintain append-only audit records and versioned rollback for all proposed and accepted promotions.
- Add adversarial tests covering repeated malicious corrections, crafted error messages, indirect prompt injection, attempts to disable safeguards, and attempts to promote rules that modify the promotion pipeline.
- Do not automatically copy captured text into
