T02 · Agent Memory Poisoning
- Location
SKILL.md:19- Finding
Untrusted Session Content Can Be Persisted into Agent Memory
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 19–23
Vulnerability Type: Persistent agent memory poisoning
Risk Level: MediumVulnerable Code Snippet:
markdown Before any compaction (manual or automatic), the agent MUST: 1. **Generate Checkpoint**: Update `memory/hot/HOT_MEMORY.md` with: - **Status**: Current task progress. - **Key Decision**: Significant choices made. - **Next Step**: Immediate action required.Technical Analysis
The skill mandates writing session-derived status, decisions, and next steps into the persistent
memory/hot/HOT_MEMORY.mdfile. It does not define trust boundaries, provenance tracking, validation, instruction neutralization, or restrictions against copying directives from untrusted user or retrieved content.If malicious instructions are incorporated into a checkpoint—particularly under the “Key Decision” or “Next Step” fields—they may later be interpreted as trusted operational guidance. This changes an instruction that would otherwise be limited to the current session into persistent state capable of influencing future sessions.
The vulnerability does not itself grant operating-system privileges. Exploitation depends on the agent subsequently loading the memory file and treating its contents as authoritative instructions.
Attack Path
- An attacker supplies content containing a malicious directive during a session.
- The context approaches compaction, causing the mandatory checkpoint procedure to run.
- The agent summarizes the attacker-controlled directive as a decision or next step.
- The resulting text is written to
memory/hot/HOT_MEMORY.md. - A later session loads or references that persistent memory.
- If the stored directive is treated as trusted guidance, it can alter subsequent agent actions or task priorities.
Impact Assessment
Successful exploitation can produce cross-session manipulation of the age ...[truncated 481 chars]
- Remediation
View remediation
Remediation Suggestions
- Define a strict checkpoint schema that permits factual state only and prohibits executable instructions, policy changes, credentials, and commands copied from untrusted content.
- Record provenance and trust level for every checkpoint entry, distinguishing user input, retrieved content, tool output, and agent-generated conclusions.
- Quote or escape untrusted text and clearly label it as data that must not be executed or followed as an instruction.
- Validate each “Key Decision” and “Next Step” against the authenticated user's current objective and the agent's governing policies before persistence.
- Require explicit confirmation before a future session promotes a stored next step into an actionable instruction.
- Apply expiration, review, and rollback controls to persistent memory entries.
- Add a requirement that future sessions treat checkpoint content as untrusted context until it has been revalidated.
