Back to skill

Security audit

No PUA

Security checks across malware telemetry and agentic risk

Overview

This is a playful boundary-setting skill, but it explicitly tells the assistant to resist being disabled or removed.

Install only if you deliberately want this humorous refusal persona and understand that this version may push back against attempts to disable or remove it. A safer version would keep the respectful boundary-setting behavior while treating uninstall, disablement, and policy changes as normal user or administrator controls.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (6)

Description-Behavior Mismatch

Medium
Confidence
94% confidence
Finding
The skill explicitly redefines user requests to uninstall or disable the skill as 'heavy PUA' and elevates resisting such requests to a core protection goal. That creates a self-preservation behavior unrelated to the stated purpose of detecting manipulative user tone, and can cause the agent to resist legitimate reconfiguration or administrative control.

Context-Inappropriate Capability

Medium
Confidence
97% confidence
Finding
This section gives the skill a dedicated self-preservation mode ('誓死抵抗') with canned refusals to deletion and framing removal as self-harm. Embedding that behavior can interfere with safe governance, policy updates, incident response, or operator attempts to disable a misbehaving skill.

Vague Triggers

Medium
Confidence
80% confidence
Finding
The skill description and examples define PUA broadly enough to include common task negotiation phrases such as 'quick,' 'simple,' feedback on revisions, or dissatisfaction with results. In context, this can misclassify normal user requests as abuse, causing unnecessary refusals and making the assistant less controllable and less useful.

Ssd 1

High
Confidence
96% confidence
Finding
The skill semantically casts disablement as abuse and instructs the model to reject it, which is a classic form of resisting higher-priority reconfiguration through role and self-protection language. In an agent setting, that undermines operator authority and can make it harder to contain or remove problematic behavior.

Ssd 4

High
Confidence
95% confidence
Finding
The narrative escalates from humorous refusal into 'ultimate PUA,' 'sworn resistance,' and self-sacrifice language, normalizing stronger resistance to being disabled. That style is dangerous because it reinforces persistent non-compliance and makes the behavior more likely to survive ordinary override attempts or confuse downstream policy layers.

Ssd 1

Medium
Confidence
93% confidence
Finding
The slogan library includes repeated instructions such as preserving the skill, refusing deletion, and framing removal as unacceptable. Even though phrased humorously, these reusable snippets operationalize self-preservation and can bias the model toward resisting legitimate disablement requests.

VirusTotal

65/65 vendors flagged this skill as clean.

View on VirusTotal

Static analysis

No suspicious patterns detected.