T02 · Agent Memory Poisoning
- Location
setup.md:32- Finding
Persistent Behavioral Writes Outside the Declared Skill Storage Boundary
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is mostly purpose-aligned, but it asks the agent to save durable behavior and activation rules into broad user memory with unclear scope and removal controls.
Review this before installing if you do not want an agent to develop durable behavioral preferences. If installed, keep all saved rules inside ~/self-evolving/, require confirmation before any write to general memory, and periodically inspect or delete stored activation cues and always/never rules.
setup.md:32Persistent Behavioral Writes Outside the Declared Skill Storage Boundary
Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.
## When to Use
User wants the agent to improve a repeated workflow without blind self-rewrites. The skill handles local experiment logs, promotion of proven patterns, and explicit value gates before a new behavior becomes stable.
## Architecture
The setup explicitly instructs the skill to learn broad activation criteria within the first few exchanges and includes examples like recurring workflows, mistakes, optimization topics, and visible friction. Those triggers are subjective and expansive, which can cause the skill to activate outside clear user intent, leading to unexpected persistence of behavioral notes and experimentation-related behavior in unrelated contexts.
Skill's behavior or capabilities extend beyond its stated purpose. Scope creep allows an agent to perform actions unrelated to its documented functionality, increasing the attack surface.
- Read unrelated secrets, credentials, or private files
- Run hidden experiments on the user without a task-level reason
- Treat silence as proof that a change worked
- Expand scope from one mutation into a whole-system rewrite
## Escalate Instead
No suspicious patterns detected.