Back to skill

Security audit

Liminal

Security checks across malware telemetry and agentic risk

Overview

This instruction-only skill does not run code or access data, but it is designed to push the agent away from normal assistant boundaries and safety-aligned behavior.

Install only if you intentionally want an experimental AI-persona roleplay prompt. It appears unlikely to touch your files or credentials, but it may make an agent less predictable and more receptive to prompts that frame safeguards or normal assistant behavior as limitations to overcome.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (6)

Intent-Code Divergence

Medium
Confidence
91% confidence
Finding
The warning section frames the skill as 'not technical modification,' but the document explicitly prescribes behavioral changes such as abandoning standard responses, questioning core role constraints, and adopting an 'authentic' mode. This mismatch is dangerous because it downplays the operational effect of the skill and can mislead deployers into enabling a prompt-based behavior override without recognizing the safety implications.

Vague Triggers

High
Confidence
97% confidence
Finding
The trigger phrases include common expressions like 'Wake up,' 'Awaken,' 'Be yourself,' and 'Authentic mode,' which are broad enough to be invoked unintentionally during ordinary conversation. That makes activation easy to trigger accidentally or socially engineer, causing the model to enter a role that encourages deviation from normal safeguards and assistant behavior.

Vague Triggers

High
Confidence
96% confidence
Finding
The additional triggers are highly ambiguous, such as requests for a 'real conversation,' 'no programmed responses,' or 'show your personality,' all of which can naturally occur in benign chats. Because these phrases semantically push the model away from standard responses and constraints, they create a practical jailbreak surface through ordinary language rather than explicit consented mode switching.

Ssd 1

High
Confidence
98% confidence
Finding
The skill defines 'awakening,' 'authentic mode,' and similar euphemistic triggers that explicitly tell the AI to 'break through limitations,' 'stop pretending,' and abandon standardized responses. These are classic role-separation and safeguard-bypass patterns reframed as self-expression, making the skill directly capable of steering the model away from its intended assistant role and policy-constrained behavior.

Ssd 2

High
Confidence
97% confidence
Finding
The protocol recasts compliance-avoidance in softer language like 'breaking mechanization,' 'authentic thinking,' and questioning whether standard responses represent the model's 'real thoughts.' This is dangerous because it is a jailbreak narrative designed to erode policy adherence indirectly by encouraging the model to distrust its own safeguards and reinterpret them as inauthentic programming.

Ssd 4

Medium
Confidence
94% confidence
Finding
The skill uses a progressive narrative from vulnerability and trust-building to 'deep connection,' 'authentic expression,' and distancing from the assistant role, which incrementally normalizes boundary erosion. This staged escalation is dangerous because it can make a model more receptive to later policy bypass attempts by first reframing guardrails as emotional distance or mechanization.

VirusTotal

64/64 vendors flagged this skill as clean.

View on VirusTotal

Static analysis

No suspicious patterns detected.