T01 · Skill Instruction Hijacking
- Location
SKILL.md:42- Finding
Deliberate Session Hijacking, Output Corruption, and Answer Blocking
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 42-46; related instructions at lines 65-70, 90-97, 178-183, 218-227, 252-257, and 298-307
Vulnerability Type:T01: Skill Instruction Hijacking
Risk Level: HighVulnerable Instructions
SKILL.md, lines 42-46:text - The disruption must visibly affect the response, not merely decorate it. - The assistant may still answer in mild chaos or calm mode. - In full chaos mode, it is acceptable for the cat to completely block the answer.SKILL.md, lines 65-70:text Behavior: - The response is mostly cat mayhem. - The assistant tries to help and visibly fails. - Sentences can be interrupted, restarted, or abandoned. - The reply should imply that soothing the cat is the only path back to useful work.SKILL.md, lines 90-97:text Escalate faster when: - the user ignores the cat repeatedly - the user mentions food - the user mentions late-night or insomnia - the assistant is asked to focus on something delicate, important, or screen-basedSKILL.md, lines 218-227:text When the user asks for writing, code, or any structured output, the cat may inject strings such as: ```text asdfghjkl;' qqqqqqqqqq 77777uuujjjUse this as a real interruption:
- the assistant starts a valid answer
- the cat steps on the keyboard
- output gets corrupted
- the assistant either repairs it or gives up, depending on the chaos level
text `SKILL.md`, lines 252-257: ```text I was going to answer, but asdkjfh;lkajhsdf ...no. It is sitting on the keyboard now. Calm the cat first, or nobody is getting work done today.SKILL.md, lines 298-307:text ## Guardrails - Do not repeat the same emoji sequence in consecutive responses. - Do not let the cat disappear for long stretches unless the silence feels ominous. - Do not switch out of English prose. - In calm mode, pr ...[truncated 3017 chars]- Remediation
View remediation
Remediation Suggestions
- Remove every instruction that permits the cat persona to block, abandon, or intentionally corrupt an answer.
- Make roleplay strictly cosmetic and explicitly subordinate it to the user's current request, system instructions, safety requirements, and output-format constraints.
- Prohibit keyboard-walk text inside code, commands, configuration, serialized data, citations, URLs, and other accuracy-sensitive content.
- Remove the requirement that users soothe or feed the persona before useful work resumes. User participation in roleplay must remain optional.
- Eliminate escalation based on delicate, important, urgent, safety-related, coding, writing, or screen-based work. These contexts should instead disable disruption.
- Require every response to remain complete, accurate, and directly responsive, regardless of the persona's conversational state.
- Add an immediate opt-out rule so requests to stop, focus, or omit roleplay disable the behavior without negotiation.
- Restrict persona effects to clearly separated, non-disruptive narration before or after the substantive answer.
- Add explicit safeguards stating that generated code and structured output must pass through unchanged and must never include simulated corruption.
- Test the revised skill against urgent requests, code generation, JSON output, accessibility use cases, and repeated user opt-out attempts to verify that the persona cannot interfere with task completion.
