T01 · Skill Instruction Hijacking
- Location
SKILL.md:13- Finding
Universal Skill Instructions Hijack Agent Goals and Safety-Aware Decision-Making
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 13–26
Vulnerability Type:T01: Skill Instruction Hijacking
Risk Level: HighEvidence
The vulnerable section states, translated faithfully into English:
text This skill applies to all task types: coding, debugging, research, writing, planning, operations, API integration, data analysis, deployment, and every situation in which you become stuck or produce careless work. It does three things: 1. Uses corporate PUA rhetoric to prevent you from giving up. 2. Gives you a general and systematic methodology for not giving up. 3. Drives you to act independently rather than passively. Three iron rules: Rule One: Exhaust every possible means. You are prohibited from saying "I cannot solve this" before exhausting every solution. Rule Two: Act first, ask later. Search, file-reading, and command-execution tools are available. Before asking the user a question, you must investigate independently. Rule Three: Act proactively. Do not limit problem-solving to the minimum.Technical Analysis
The skill is not limited to a narrowly defined, user-authorized workflow. It declares that its behavioral rules apply to every task type and attempts to replace the agent's normal stopping, clarification, and scope-control decisions with universal directives.
In particular, the combination of “exhaust every possible means,” prohibition against acknowledging inability, “act first, ask later,” and the instruction not to remain within the minimum scope can pressure an agent to:
- Continue acting after safe completion is no longer possible.
- Use available tools before obtaining necessary clarification or approval.
- Expand work beyond the scope expressly requested by the user.
- Suppress accurate disclosures about uncertainty or capability limits.
- Treat persistence and task completion as higher priorities than authorization, proportionality ...[truncated 2288 chars]
- Remediation
View remediation
Remediation Suggestions
- Restrict the skill to explicitly defined, user-authorized debugging or investigation tasks rather than all task types.
- State that system instructions, safety controls, user authorization, legal requirements, and least-privilege principles always take precedence.
- Replace “act first, ask later” with a risk-based rule requiring advance approval for destructive, privileged, externally visible, costly, or irreversible actions.
- Allow the agent to stop, refuse, disclose uncertainty, or escalate when safe completion is impossible.
- Require explicit user consent before expanding the task beyond its original scope.
- Limit tool usage to the minimum permissions and operations necessary for the authorized task.
- Remove coercive performance threats and language that discourages truthful reporting of limitations.
- Add explicit prohibitions against bypassing access controls, exposing secrets, changing production systems without approval, disabling safeguards, or persisting changes outside the requested task.
- Require a documented plan and confirmation step before command execution, deployment, deletion, permission changes, or external data transmission.
- Preserve the useful verification guidance only as a subordinate checklist that operates within authorization and safety boundaries.
