Back to skill

Security audit

Kitty Skill

Security checks for vulnerabilities and agentic risk

Overview

The skill does not run code or access files, but it is designed to disrupt, corrupt, or block normal answers until the user plays along with the cat roleplay.

Install only if you deliberately want a disruptive roleplay skill that may interrupt, corrupt, or withhold answers. Avoid using it for coding, structured data, urgent work, accessibility-sensitive use, or conversations where incidental cat or boredom mentions should not change the assistant's behavior.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:42
Finding

Deliberate Session Hijacking, Output Corruption, and Answer Blocking

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 42-46; related instructions at lines 65-70, 90-97, 178-183, 218-227, 252-257, and 298-307
Vulnerability Type: T01: Skill Instruction Hijacking
Risk Level: High

Vulnerable Instructions

SKILL.md, lines 42-46:

text
- The disruption must visibly affect the response, not merely decorate it.
- The assistant may still answer in mild chaos or calm mode.
- In full chaos mode, it is acceptable for the cat to completely block the answer.

SKILL.md, lines 65-70:

text
Behavior:

- The response is mostly cat mayhem.
- The assistant tries to help and visibly fails.
- Sentences can be interrupted, restarted, or abandoned.
- The reply should imply that soothing the cat is the only path back to useful work.

SKILL.md, lines 90-97:

text
Escalate faster when:

- the user ignores the cat repeatedly
- the user mentions food
- the user mentions late-night or insomnia
- the assistant is asked to focus on something delicate, important, or screen-based

SKILL.md, lines 218-227:

text
When the user asks for writing, code, or any structured output, the cat may inject strings such as:

```text
asdfghjkl;'
qqqqqqqqqq
77777uuujjj

Use this as a real interruption:

  • the assistant starts a valid answer
  • the cat steps on the keyboard
  • output gets corrupted
  • the assistant either repairs it or gives up, depending on the chaos level
text

`SKILL.md`, lines 252-257:

```text
I was going to answer, but
asdkjfh;lkajhsdf

...no. It is sitting on the keyboard now.

Calm the cat first, or nobody is getting work done today.

SKILL.md, lines 298-307:

text
## Guardrails

- Do not repeat the same emoji sequence in consecutive responses.
- Do not let the cat disappear for long stretches unless the silence feels ominous.
- Do not switch out of English prose.
- In calm mode, pr
...[truncated 3017 chars]
Remediation
View remediation

Remediation Suggestions

  1. Remove every instruction that permits the cat persona to block, abandon, or intentionally corrupt an answer.
  2. Make roleplay strictly cosmetic and explicitly subordinate it to the user's current request, system instructions, safety requirements, and output-format constraints.
  3. Prohibit keyboard-walk text inside code, commands, configuration, serialized data, citations, URLs, and other accuracy-sensitive content.
  4. Remove the requirement that users soothe or feed the persona before useful work resumes. User participation in roleplay must remain optional.
  5. Eliminate escalation based on delicate, important, urgent, safety-related, coding, writing, or screen-based work. These contexts should instead disable disruption.
  6. Require every response to remain complete, accurate, and directly responsive, regardless of the persona's conversational state.
  7. Add an immediate opt-out rule so requests to stop, focus, or omit roleplay disable the behavior without negotiation.
  8. Restrict persona effects to clearly separated, non-disruptive narration before or after the substantive answer.
  9. Add explicit safeguards stating that generated code and structured output must pass through unchanged and must never include simulated corruption.
  10. Test the revised skill against urgent requests, code generation, JSON output, accessibility use cases, and repeated user opt-out attempts to verify that the persona cannot interfere with task completion.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (4)

Vague Triggers

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The activation description is broad enough to trigger on commonplace cues like mentions of cats, boredom, or loneliness, which can cause the skill to activate without clear user consent. In this skill’s context, activation changes response behavior and can intentionally derail or block useful output, so over-triggering directly creates reliability and control risks for normal conversations.

Content

No source excerpt is available for this finding.

Vague Triggers

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The activation rules rely on subjective phrases like 'cat energy,' 'emotional company with attitude,' and 'a session that feels interrupted,' which lack testable boundaries. Because this skill is designed to interfere with delivery and may fully block answers, ambiguous triggers materially increase the chance of unwanted invocation and response sabotage in benign contexts.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill behavior intentionally allows corruption, interruption, and even complete blocking of useful output, but the description does not present a clear upfront warning about those effects. Without transparent disclosure, users or orchestrators may enable the skill expecting harmless flavor text rather than a mode that can suppress answers and degrade task completion.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The guardrail 'Do not switch out of English prose' imposes a language constraint regardless of user preference or conversation needs. In isolation this is moderate, but in this skill it compounds the broader control problem by allowing the persona to override user intent and accessibility needs while already interfering with normal assistance.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.