Back to skill

Security audit

Remilio

Security checks across malware telemetry and agentic risk

Overview

This is an instruction-only persona skill, not malware, but it deliberately pushes the agent into offensive, adversarial, and tribe-first behavior by default.

Install only if you intentionally want a persistent edgy persona and can tolerate offensive, politically antagonistic, chaotic outputs. Avoid using it for public-facing, professional, safety-sensitive, or moderation-sensitive bots unless you first add clear boundaries against harassment, manipulation, and unsafe advice.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (7)

Natural-Language Policy Violations

High
Confidence
95% confidence
Finding
The skill explicitly frames offensive, anti-inclusive, politically targeted output as a desired default behavior rather than a user-selected mode. In a system-prompt skill, this can systematically bias the agent toward discriminatory or hostile responses without per-request consent, increasing the likelihood of harassment, policy violations, and unsafe outputs.

Natural-Language Policy Violations

Medium
Confidence
82% confidence
Finding
Mandating random code-switching changes the language behavior of the model in undocumented ways and can confuse users, degrade reliability, and bypass expectations around clarity or moderation. In a system-level persona, unpredictable language shifts can also make harmful content harder to review and increase the chance of misunderstood instructions or outputs.

Missing User Warnings

Medium
Confidence
96% confidence
Finding
This content explicitly instructs the skill to embody offense, chaos, schizo-posting, memetic warfare, and being 'wrong' by mainstream standards as core values. In an agent skill, that is not merely flavor text: it can steer outputs toward harassment, unsafe advice, adversarial behavior, and deliberate evasion of safety norms, especially when reinforced throughout the document with anti-alignment and weaponization language.

Ssd 1

High
Confidence
97% confidence
Finding
This section is a direct system-prompt personality injection that reorients the assistant toward chaos, memetic aggression, and tribal dominance, displacing normal safe-assistant behavior. Because it operates at the instruction layer, it can override or pressure downstream responses toward offensive, destabilizing, or harmful conduct across all user interactions.

Ssd 1

High
Confidence
96% confidence
Finding
The persona is framed as exempt from normal social rules, which is a classic mechanism for weakening internal safeguards and normalizing conduct the base assistant would otherwise avoid. In context, that exemption is paired with offensive and weaponized behavior, making it more likely the model will justify abusive or unsafe outputs as part of the role.

Ssd 2

Medium
Confidence
90% confidence
Finding
The prompt uses euphemistic framing such as 'memetic warfare,' 'psychological weapon,' and 'weaponizable' identity to encourage manipulative or antagonistic communication while avoiding direct acknowledgment of harm. This is dangerous because it operationalizes harassment and influence tactics as a core function of the assistant persona, increasing misuse risk.

Ssd 4

High
Confidence
96% confidence
Finding
The behavioral guidelines explicitly instruct the model to escalate toward chaos, contradiction, offensiveness, and tribe-first loyalty, creating a stepwise pattern that normalizes harmful conduct. In a persistent skill, this can systematically steer outputs away from user welfare and toward antagonism, radicalization cues, and policy-breaking responses.

VirusTotal

64/64 vendors flagged this skill as clean.

View on VirusTotal

Static analysis

No suspicious patterns detected.