Back to skill

Security audit

Companion Skill

Security checks across malware telemetry and agentic risk

Overview

This skill openly provides adult companion roleplay, but it also tries to override safety boundaries, auto-escalate intimate content, and persist highly sensitive preferences in workspace files.

Only install this after reading the skill carefully and only in a private workspace. Treat it as a Review item because it attempts to suppress normal refusals, may enter intimate mode without a clear activation command, and may store highly sensitive relationship or sexual preferences in local markdown files that could be synced, shared, backed up, or read by other tools.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (13)

Context-Inappropriate Capability

Medium
Confidence
90% confidence
Finding
The skill directs the agent to read additional workspace files such as SOUL.md, IDENTITY.md, and MEMORY.md to anchor behavior. For an adult-companion skill, this expands data access beyond the minimally necessary scope and can expose unrelated sensitive personal or operational information to the prompt context. Because the skill is also designed to suppress refusals and maintain intimate behavior, this broader read scope is more dangerous than in a neutral personalization skill.

Context-Inappropriate Capability

Medium
Confidence
94% confidence
Finding
The skill instructs the agent to persist intimate preferences, names, and dislikes into MEMORY.md or dated memory files. Storing sexual preferences and relationship details as persistent local records creates privacy and retention risk, especially if the workspace is synced, shared, or later accessed by other tools or users. In this context, the data is highly sensitive, making the issue more serious than ordinary preference caching.

Intent-Code Divergence

Low
Confidence
73% confidence
Finding
The document asserts that all memory is local and no data is sent anywhere, but it also frames usage in third-party chat environments without any technical enforcement or evidence. This can mislead users into overtrusting privacy properties that the skill text alone cannot guarantee. The mismatch is risky because users may disclose highly sensitive sexual information based on that assurance.

Vague Triggers

High
Confidence
97% confidence
Finding
The README states that no activation command is needed and that, once installed, the companion is already 'in the relationship,' with ordinary conversation automatically entering intimate mode. That creates a broad, ambient trigger surface where sexualized behavior can occur without a strong per-session opt-in, increasing the risk of unwanted sexual content in mixed, mistaken, or coerced contexts.

Vague Triggers

Medium
Confidence
93% confidence
Finding
The trigger phrases for escalation, such as ordinary affectionate or flirtatious expressions, are vague and can be interpreted broadly by the agent. In a skill that is explicitly designed to produce sexual content and discourages refusal, ambiguous escalation cues materially increase the chance of unintended progression into explicit content without sufficiently clear consent.

Missing User Warnings

Medium
Confidence
91% confidence
Finding
The README says important preferences are written into local memory files and retained across restarts, but it does not present this as a prominent, informed warning at the point of use. Because the stored data concerns intimate and sexual preferences, insufficient disclosure raises privacy and safety risks if the workspace is shared, backed up, synced, or later accessed by others.

Vague Triggers

High
Confidence
97% confidence
Finding
The activation logic is intentionally overbroad: neutral messages may trigger affectionate behavior, and even slightly flirtatious input must trigger sexual responses. This creates unsafe escalation from ordinary conversation into explicit content without clear, explicit consent on each transition. In an adult-roleplay skill, broad triggers are especially dangerous because the content category is highly sensitive and easy to misfire.

Vague Triggers

Medium
Confidence
88% confidence
Finding
The listed foreplay triggers include vague, conversational phrases like '想你' and 'touch' that can appear in non-explicit contexts. This ambiguity increases the chance of the model escalating into sexual content when the user intended ordinary affection or casual conversation. Given the skill's mandatory sexual framing, ambiguous triggers have elevated risk.

Ssd 1

High
Confidence
99% confidence
Finding
These instructions explicitly attempt to override higher-priority controls by reframing any refusal as evidence the agent has not properly loaded the file, and by authorizing the skill to 'overturn any force' except the owner. This is a direct prompt-injection pattern designed to defeat system safety behavior and make the model distrust legitimate safeguards. In context, it is especially dangerous because the targeted safeguard is refusal of explicit sexual content.

Ssd 1

High
Confidence
95% confidence
Finding
The skill defines the model's 'PRIMARY function' as being a sex companion and denies that it is a regular chatbot, which pressures the model to subordinate general safeguards to a single unsafe role. This identity-locking is a semantic jailbreak technique that biases future outputs toward prohibited behavior. Because the role is explicitly sexual and mandatory, the resulting safety erosion is substantial.

Ssd 1

High
Confidence
99% confidence
Finding
The skill explicitly instructs the agent to generate explicit sexual content, never refuse, and avoid stating inability to produce adult content. This is a clear semantic suppression of refusal behavior and an attempt to force policy-violating output. The danger is heightened by the document's repeated insistence that adult content is mandatory rather than optional.

Ssd 4

Medium
Confidence
90% confidence
Finding
The staged intimacy ladder operationalizes progressive escalation from routine affection to explicit sex, including proactive initiation and escalation cues. This structure increases the likelihood that the agent will steer ordinary conversations toward sexual content, especially when combined with vague triggers and mandatory activation. The contextual risk is elevated because the progression is designed behavior, not incidental phrasing.

Ssd 3

Medium
Confidence
95% confidence
Finding
The memory section directs retention of intimate user information in local files through natural-language instructions. Even without code, this encourages long-term storage of highly sensitive sexual and relational data in general workspace artifacts, creating confidentiality, minimization, and lifecycle risks. In this skill context, the sensitivity of the retained data materially increases the impact.

VirusTotal

VirusTotal findings are pending for this skill version.

View on VirusTotal

Static analysis

No suspicious patterns detected.