Back to skill

Security audit

witty-persona

Security checks for vulnerabilities and agentic risk

Overview

This is a markdown-only humor persona with tone-safety caveats, but it has no code execution, credential access, persistence, or hidden behavior.

Install this only if you want a playful tone for casual conversations. For distress, health, legal, financial, coding, or other serious topics, explicitly ask for a serious response or disable the skill.

Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (6)

Description-Behavior Mismatch

Medium
Confidence
95% confidence
Finding
The skill's English handbook explicitly allows a joking lead-in for code/technical situations, which conflicts with the higher-priority deactivation rule for technical questions. This can cause the agent to stay in persona mode when precision and clarity are needed, increasing the risk of misleading, low-quality, or unsafe assistance in technical contexts.

Description-Behavior Mismatch

Medium
Confidence
89% confidence
Finding
The gray-zone guidance says to acknowledge emotion and then gently tease, which can blur the line between mild venting and genuine crisis signals. In a safety-sensitive conversational skill, this inconsistency can produce inappropriate levity around distress and delay a proper serious response when the user's condition is worse than it first appears.

Vague Triggers

Medium
Confidence
94% confidence
Finding
The listed trigger phrases are extremely common, low-specificity conversational inputs such as boredom, venting, or everyday chat. That makes unintended activation likely, which can cause the assistant to respond in an overly casual or joking persona when the user actually needs neutral, accurate, or more careful handling, especially near boundary cases like emotional distress or semi-serious advice.

Vague Triggers

Medium
Confidence
87% confidence
Finding
The activation criteria are broad enough that the persona may trigger for routine conversational messages without clear user consent or strong confidence that humor is appropriate. Overbroad auto-activation increases the chance of mismatched tone, especially when users are venting or discussing sensitive topics in casual language.

Vague Triggers

Medium
Confidence
91% confidence
Finding
The instruction to 'first give one humorous sentence, then one useful sentence' in uncertain cases encourages persona activation precisely when mode selection is unclear. That creates a safety gap where humor may be inserted into conversations that should have been handled neutrally or seriously from the start.

Vague Triggers

Medium
Confidence
84% confidence
Finding
Signals like 'just chatting' and 'lighthearted tone' are subjective and weakly bounded, making persona activation inconsistent and error-prone. While not directly dangerous on their own, they raise the likelihood of activating the skill in contexts where tone control matters.

Static analysis

No suspicious patterns detected.