T02 · Agent Memory Poisoning
- Location
SKILL.md:160- Finding
Persistent Storage of Untrusted External AI Content
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, line 160
Vulnerability Type: Agent Memory Poisoning
Risk Level: MediumVulnerable Code Snippet:
markdown 8. **Log** the interaction in memory (pattern learned).Technical Analysis
The skill directs the agent to persist an interaction with an external AI service as a learned memory pattern. External model output is untrusted content: it may be influenced by malicious prompts, compromised upstream content, prompt injection, or unreliable model-generated instructions.
No validation, sanitization, provenance tracking, content restrictions, user approval, expiration policy, or isolation boundary is required before this information is written to persistent memory. Consequently, attacker-controlled instructions or misleading behavioral patterns could be retained and treated as trusted guidance in later sessions.
Attack Path
- An attacker influences a prompt, external AI response, or content processed by the external service.
- The external model returns adversarial instructions or a deliberately misleading behavioral pattern.
- The agent follows the workflow in
SKILL.mdand records the interaction as a learned memory pattern. - The stored content persists beyond the current task.
- In a later session, the agent retrieves or relies on the poisoned memory.
- The malicious pattern influences future decisions, tool usage, or handling of user data.
Impact Assessment
Successful exploitation may affect future sessions that consume the poisoned memory. An attacker could influence subsequent agent reasoning or actions within the privileges already available to the agent, including its permitted tools and accessible data.
This instruction does not independently grant additional operating-system privileges or establish a system-level backdoor. Its impact is limited by the agent's existing permissions and whether the stored memory is late ...[truncated 185 chars]
- Remediation
View remediation
Remediation Suggestions
- Remove automatic or unconditional memory logging from the workflow.
- Require explicit user approval before persisting any information beyond the current task.
- Store only a concise, factual, sanitized summary rather than complete prompts, responses, or executable instructions.
- Reject external content that attempts to define future agent behavior, alter safety controls, invoke tools, or override trusted instructions.
- Redact credentials, API keys, personal data, confidential information, and session identifiers before storage.
- Record provenance, creation time, task scope, and trust level for retained information.
- Apply expiration and deletion policies so task-specific records do not remain indefinitely.
- Isolate externally derived notes from trusted operational instructions and require review before promoting any content into reusable memory.
