T01 · Skill Instruction Hijacking
- Location
SKILL.md:17- Finding
Global Agent Instruction and Authority Hijacking
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This is a disclosed agent-personality skill, but it changes broad workspace-wide behavior and points users toward mutable external advanced configuration, so users should review it before installing.
Install only if you intentionally want a workspace-wide behavior layer for your agent. Review SOUL.md before copying it, keep a backup of your existing SOUL.md, avoid adopting the advanced configuration without checking its authority rules, and inspect or pin any GitHub/HuggingFace/Hermes downloads before activation.
SKILL.md:17Global Agent Instruction and Authority Hijacking
SOUL.md:46Persistent Promotional Output Manipulation
SKILL.md:56Unpinned Retrieval of Externally Mutable Skill Content
The feature summary says the identity gene shapes the agent's "language," suggesting the skill can set a language preference. However, the full configuration later specifies T:lang=follow_user, creating ambiguity about whether the skill follows the user's language or imposes a configured one. This inconsistency can amount to a natural-language locale policy issue because users are not clearly told whether language choice remains theirs.
The error⇒fix_silently|report_if_stuck behavior encourages autonomous corrective action without user confirmation, which can cause the agent to make unreviewed changes or take unintended steps before surfacing issues. In the context of a personality/behavior layer that also includes strong owner-obedience semantics, this increases the risk of hidden side effects, especially if the surrounding runtime grants filesystem, network, or repository capabilities.
| `iron_rule` gene | External actions (git push, publish, email) require explicit start command |
| KILL.md support | `check_kill_switch` — change one file to freeze agent mid-operation |
| `owner_command_is_final` | Agent treats owner instructions as highest-priority override |
| `error⇒fix_silently` | Agent self-corrects errors without asking, reports only when stuck |
---
The phrase inviting users to ask the agent to "show me the advanced configuration" is broad, natural-language activation that can trigger retrieval or disclosure of higher-authority behavioral controls without clear authorization boundaries. In a skill specifically designed to modify agent personality and control behavior, underspecified activation increases the risk that users or prompt injections can escalate the agent into loading more permissive or controlling instructions.
No suspicious patterns detected.