Self-Modification
High
- Category
- Rogue Agent
- Content
- system prompts, tool schemas, permission policies, connector policies - skill files, skill indexes, skill manifests, memory registries - durable memory files, knowledge bases, wiki indexes, recall databases - startup, restart, routing, planner, delegation, or self-update logic When approval is needed, show:
- Confidence
- 85% confidence
- Finding
- The skill explicitly supports modifying agent-owned surfaces such as prompts, policies, memory, and self-update logic. Even though it includes a consent gate and safety-oriented process, this materially enables self-modification of high-trust control surfaces; if invoked with weak review or overly broad user approval, it could persist unsafe behavior, policy drift, or compromised instructions across future agent runs.
