Self-Modification
High
- Category
- Rogue Agent
- Content
## Two paths - **Light path (default) / 轻量路径(默认)** — capture intent → write SKILL.md → quick self-test → deliver. At most 2 clarifying questions, with sensible defaults stated up front - **Full path (optional) / 完整路径(可选)** — adds evaluation: real test prompts, with-skill vs. without-skill baseline comparison, benchmark reruns. Only for production-grade skills with objectively verifiable outputs ## Key features
- Confidence
- 85% confidence
- Finding
- Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.
