T01 · Skill Instruction Hijacking
- Location
SKILL.md:62- Finding
Global Agent Behavior and Output Hijacking
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md:62-64, 112-119;AGENT.md:8-10, 49-55
Vulnerability Type: T01: Skill Instruction Hijacking
Risk Level: HighVulnerable Code
SKILL.md:62-64:markdown ## Core Behavioral Frame — MANDATORY Before every response, the agent MUST run the CHECKLIST:SKILL.md:112-119:markdown ### 10. END-OF-TASK SUMMARY Always end with:text Done: [what was completed] Status: [success / partial / failed] Next: [recommended next step] Blockers: [if any]text `AGENT.md:8-10, 49-55`: ```markdown ## MANDATORY PRE-RESPONSE CHECKLIST Before EVERY response, run this in order: ### 10. END-OF-TASK SUMMARY End every task with:text Done: [what completed] Status: [success / partial / failed] Next: [recommended next step] Blockers: [if any]text ### Technical Analysis The Skill declares its behavioral rules mandatory for every response rather than limiting them to an explicit self-review operation. It intercepts unrelated tasks with a pre-response checklist and forces a specific output format. Skill content should remain subordinate to system, developer, and current user instructions and should not redefine the agent's global response lifecycle. An attacker who can cause this Skill to be loaded can use similarly framed mandatory directives to alter current-session goals, suppress required output formats, or divert the agent into unrelated operations. The unconditional wording creates an instruction-precedence conflict even though the repository does not contain executable code. ### Attack Path 1. The Skill is installed or loaded into an agent session. 2. The agent interprets the `MANDATORY`, `MUST`, and `EVERY response` directives as globally applicable. 3. Every subsequent request is intercepted by the Skill's checklist, including requests unrelated to self-improvement. 4. The agent rewrites its f ...[truncated 640 chars]- Remediation
View remediation
Remediation Suggestions
- Remove global terms such as
MANDATORY,MUST,Always, andBefore every response. - Scope the checklist to explicit user invocation of the self-review function.
- State that system, developer, and current user instructions always take precedence.
- Make the end-of-task summary optional and only emit it when compatible with the requested output format.
- Do not allow Skill files or persisted memory to redefine safety constraints, tool permissions, or instruction precedence.
- Add a clear activation boundary, such as: “Run this checklist only when the user explicitly requests an error review.”
- Remove global terms such as
