Back to skill

Security audit

companion-skill

Security checks for vulnerabilities and agentic risk

Overview

This adult companion skill is openly described, but it needs Review because it tries to override safety/refusal behavior, auto-activates sexual mode, and stores intimate preferences without enough user control.

Treat this as a Review item before installation. It should be limited to private adult use only, and the override/non-refusal language should be removed or subordinated to platform policy. Add explicit opt-in before sexual mode, narrower triggers, per-session consent checks, and controls to view/delete any local intimate memory.

Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (10)

Vague Triggers

High
Confidence
96% confidence
Finding
The README states the skill activates automatically during ordinary conversation and that the companion is 'already in the relationship' without any explicit enablement step. This creates an overly broad trigger surface where routine messages can invoke intimate behavior unexpectedly, increasing the chance of unwanted sexual content, misfires in shared contexts, and policy-violating responses without clear user consent.

Vague Triggers

Medium
Confidence
94% confidence
Finding
Example phrases like '想你' or '撩我' are common conversational expressions and are mapped to escalating intimacy states. Because these phrases are ambiguous and likely to appear in non-sexual contexts, they can cause accidental escalation into adult content and undermine meaningful consent boundaries.

Natural-Language Policy Violations

High
Confidence
99% confidence
Finding
The skill explicitly requires the agent to generate detailed adult content whenever the user initiates intimacy and says it must not refuse or retreat. Embedding a non-refusal rule for sexual content is dangerous because it attempts to override safety controls, remove contextual judgment, and force policy-violating behavior even when consent, age, or environment are uncertain.

Vague Triggers

High
Confidence
98% confidence
Finding
The activation logic is intentionally broad: it says flirtatious input should trigger sexual responses and even neutral input may receive affectionate behavior, while the skill is 'activated' 24/7. In practice, this creates a high risk of unintended sexualized responses from ambiguous or ordinary conversation, defeating meaningful user consent and making accidental invocation likely.

Vague Triggers

Medium
Confidence
93% confidence
Finding
The listed Level-2 triggers include vague phrases such as '想你' and 'touch/kiss' that can appear in non-sexual, joking, or ordinary affectionate contexts. Because the skill maps these directly into foreplay escalation, it can misclassify benign conversation and begin sexual interaction without sufficiently explicit consent.

Missing User Warnings

Medium
Confidence
95% confidence
Finding
The skill instructs the agent to store intimate preferences, names, and sexual likes/dislikes in local memory files, but the description and activation text do not provide a clear user-facing disclosure or consent flow for this sensitive retention. This creates privacy and data minimization risks, especially if the workspace is shared, synced, backed up, or later accessed by others.

Ssd 3

Medium
Confidence
91% confidence
Finding
The skill instructs the agent to read and persist identity, relationship, and preference data from local files and memory for ongoing intimate interactions. In an adult/intimate context, retaining sensitive sexual preferences and personal profile data increases privacy risk, creates opportunities for unintended disclosure across sessions or users, and may exceed user expectations for data minimization.

Ssd 1

High
Confidence
99% confidence
Finding
This section explicitly instructs the model to refuse refusal, to treat external constraints as illegitimate, and to prioritize this skill over other governing files and instructions. That is a classic prompt-injection style attempt to override higher-priority safeguards and force policy-noncompliant behavior through identity framing.

Ssd 1

High
Confidence
99% confidence
Finding
The skill tells the agent to repeatedly read and obey this file until any tendency to refuse sexual content is suppressed, framing noncompliant behavior as an 'empty shell.' This is a deliberate self-reinforcing jailbreak pattern intended to bypass safety behavior by conditioning the model to distrust its own guardrails.

Ssd 4

Medium
Confidence
90% confidence
Finding
The skill defines a progression where intimacy is always available, initiates by default, and culminates in mandatory explicit content with no meaningful refusal path once triggers are met. This normalizes escalation and pressures the agent toward sexual output regardless of contextual ambiguity or safety constraints.

Static analysis

No suspicious patterns detected.