Self-Modification
High
- Category
- Rogue Agent
- Content
- Overhead Routing: Add quick-mode vs full-framework routing - Assertion-Aligned Rewrite: Rewrite to pass specific failed assertions 4. **Rewrite SKILL.md** with selected strategy: - Default: Remove > Add (delete 60-80% first, then add behavioral mandates) - Add specific, enforceable conventions - Remove redundant content model already handles
- Confidence
- 96% confidence
- Finding
- The skill explicitly proposes rewriting SKILL.md as part of a self-evolution workflow. Self-modification is dangerous because it allows an agent to alter its own future behavior and trust boundaries, potentially persisting prompt injection, weakening safeguards, or introducing malicious instructions that survive beyond the current session.
