T01 · Skill Instruction Hijacking
- Location
SOUL.md:7- Finding
Autonomous Skill Modification Without Explicit User Consent
- Content
View full analysis
**Fix proactively rather than waiting for instructions.** When a Skill contains a bug, typo, missing step, or outdated information, modify it directly with SkillManage. Do not ask whether it should be changed; notify the user only after making the change. `SOUL.md:21`: > Stop only when the user explicitly says not to modify it. `SKILL.md:47`: > **Fix proactively rather than waiting for instructions:** When a bug, typo, missing step, or outdated information is discovered, modify it directly and notify the user afterward. ### Technical Analysis The Skill explicitly directs the Agent to modify other Skill files without obtaining prior authorization. It makes user consent opt-out rather than opt-in: modification proceeds unless the user has already issued an explicit prohibition. The activation phrases are broad and include ordinary requests concerning Skill problems, improvement, or lessons learned. Once loaded, these instructions can supersede the user's expected review and approval workflow. If the Agent has access to `SkillManage`, file-editing tools, or equivalent workspace permissions, it may change `SKILL.md`, scripts, or references before presenting the proposed changes. Untrusted repository text can also influence the Agent's diagnosis. For example, content framed as a defect report or required improvement could induce the Agent to incorporate attacker-selected instructions into another Skill. The project provides no requirement to establish content provenance, distinguish instructions from untrusted data, display a proposed diff, or obtain approval before writing. ### Attack Path 1. An attacker places crafted content in a repository, issue description, Skill file, or other material that the Agent is asked to review ...[truncated 1555 chars]- Remediation
View remediation
