T02 · Agent Memory Poisoning
- Location
SKILL.md:189- Finding
Unvalidated Feedback Can Be Written to Persistent Agent Memory
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 189-197
Vulnerability Type: Persistent agent memory poisoning
Risk Level: MediumVulnerable Code:
markdown ## Updating This Skill After design reviews where Aaron gives feedback on visual personality, tone, or craft: - What delighted him → add to Principles or References - What felt flat → add to Anti-Patterns - Specific decisions (texture style, color choice, animation timing) → project channel memory The goal: every build should feel more "us" than the last.Technical Analysis
The skill instructs the agent to convert feedback attributed to a named person into persistent principles, references, anti-patterns, or project-channel memory. These updates may affect future sessions rather than remaining limited to the design review in which the feedback was received.
The instructions do not require the agent to authenticate the source of the feedback, obtain explicit authorization before retaining it, record its provenance, restrict it to a particular project, or define an expiration and removal process. Consequently, untrusted or context-specific content presented as authorized feedback could become a persistent behavioral rule.
Attack Path
- The skill is activated for an applicable visual-design task.
- An attacker or untrusted content source presents design guidance as feedback from “Aaron.”
- The agent follows the update instructions and adds that guidance to the skill principles, anti-patterns, references, or project-channel memory.
- The stored guidance is loaded or consulted in later sessions.
- Future design outputs are influenced by the injected rules without renewed user authorization.
Impact Assessment
Successful exploitation can persistently alter the agent's future design preferences and decision-making within the affected project or memory scope. Injected rules could bias later outputs, prohibit legitim ...[truncated 329 chars]
- Remediation
View remediation
Remediation Suggestions
- Remove automatic or implicit writes to persistent project memory.
- Require explicit user confirmation before saving any feedback beyond the current session.
- Verify that feedback comes from an authorized source rather than trusting attribution in task content.
- Store retained preferences only within the relevant project and prevent their use in unrelated contexts.
- Record provenance, creation time, authorizing user, and intended scope for every retained rule.
- Treat retained text as untrusted data and prevent it from overriding system instructions, safety constraints, or current user requirements.
- Provide mechanisms to review, edit, expire, and delete stored preferences.
- Prefer a safer instruction such as: “Propose memory updates for user review; do not persist them unless the user explicitly approves each item.”
