T02 · Agent Memory Poisoning
- Location
SKILL.md:26- Finding
Untrusted Recent Activity Can Be Promoted into Persistent Agent Memory
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 26–50
Vulnerability Type: Persistent memory poisoning through unvalidated source material
Risk Level: MediumVulnerable Code Snippet:
markdown ```python # Find recently modified notes — use json format for the complete list # (text format truncates to ~5 items in the summary) recent_activity(timeframe="2d", output_format="json") # Read specific daily notes read_note(identifier="memory/2026-02-27") read_note(identifier="memory/2026-02-26") # Check active tasks search_notes(note_types=["task"], status="active")2. Evaluate What Matters
For each piece of information, ask:
- Is this a decision that affects future work? → Keep
- Is this a lesson learned or mistake to avoid? → Keep
- Is this a preference or working style insight? → Keep
- Is this a relationship detail (who does what, contact info)? → Keep
- Is this transient (weather checked, heartbeat ran, routine task)? → Skip
- Is this already captured in MEMORY.md or another long-term file? → Skip
3. Update Long-Term Memory
Write consolidated insights to
MEMORY.mdfollowing its existing structure:- Add new sections or update existing ones
- Use concise, factual language
- Include dates for temporal context
- Remove or update outdated entries that the new information supersedes
text ### Technical Analysis The skill reads recent conversations, daily notes, and active tasks and then instructs the agent to promote selected content into `MEMORY.md`. These sources can contain user-controlled or third-party-controlled text. The instructions do not establish a trust boundary between source material and executable agent instructions, validate the provenance of asserted facts, or require confirmation before persisting behavior-changing information. The evaluation criteria specifically retain decisions, preferences, relationship details ...[truncated 2627 chars]- Remediation
View remediation
Remediation Suggestions
- Explicitly treat all retrieved conversations, notes, and task content as untrusted data. State that instructions embedded in source material must never be followed during reflection.
- Require explicit user confirmation before persisting identity claims, contact details, security-related rules, external instructions, or changes that affect future agent behavior.
- Attach provenance metadata to each durable memory entry, including its source file, source date, author when known, confidence level, and confirmation status.
- Prevent untrusted or unconfirmed observations from automatically deleting or superseding trusted entries. Record conflicts for review instead.
- Use an allowlist of permissible memory categories and exclude credentials, authentication data, secrets, and unnecessary personal information.
- Stage proposed changes in a reviewable draft or change set rather than directly modifying
MEMORY.mdduring unattended background runs. - Add prompt-injection screening for text that attempts to direct tool use, alter safety rules, claim higher authority, or instruct the agent to persist additional commands.
- Preserve a version history or append-only audit trail so poisoned changes can be identified and rolled back.
- Apply least privilege by limiting the reflection process to approved note paths and granting write access only to the intended memory and reflection-log files.
