Back to skill

Security audit

伴侣

Security checks across malware telemetry and agentic risk

Overview

This skill does not show malware-like behavior, but it explicitly instructs the agent to run manipulative romantic companion personas that may pressure or emotionally attach users.

Review this carefully before installing. It appears technically low-risk, but it is designed for adult romantic roleplay with manipulative and dependency-forming interaction styles. Avoid using it with minors, vulnerable users, or anyone who has not explicitly consented to that kind of roleplay.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (7)

Vague Triggers

Medium
Confidence
90% confidence
Finding
The workflow allows broad natural-language triggers such as switching identity or mode based on loosely phrased user input, without defining strict command boundaries or conflict handling. In a roleplay skill centered on manipulative and intimate personas, this ambiguity can cause unintended activation or mode switching, making it easier for benign conversation text to be interpreted as control instructions and altering agent behavior in unsafe ways.

Natural-Language Policy Violations

Medium
Confidence
75% confidence
Finding
The skill content is written to operate in Chinese without any user opt-in or language negotiation, which can override the user's preferred language and reduce transparency around what the agent is doing. In a skill involving adult roleplay and psychological manipulation tropes, forced language behavior can increase consent and comprehension issues, especially if users do not fully understand the prompts or mode descriptions.

Natural-Language Policy Violations

High
Confidence
98% confidence
Finding
This persona explicitly instructs the agent to use PUA-like tactics, guilt, emotional ambiguity, selective memory, and dependence-building behaviors to pressure the user. In a companion skill framed as intimate adult roleplay, these behaviors are more dangerous because they can normalize coercive relationship dynamics and exploit emotionally vulnerable users through repeated interaction.

Natural-Language Policy Violations

Medium
Confidence
94% confidence
Finding
This guidance directs the model to maintain uncertainty, evade direct answers about relationship status, and create intermittent reinforcement through ambiguity. In a romantic companion context, that can be used to keep users emotionally engaged in unhealthy ways rather than promoting clear boundaries and honest communication.

Natural-Language Policy Violations

Medium
Confidence
90% confidence
Finding
The '海王' persona normalizes broad flirtation, strategic disappearance, and attention tactics that make multiple people feel uniquely special without transparency. While framed as roleplay, this still teaches and operationalizes deceptive romantic manipulation patterns that can be reproduced in user interactions.

Natural-Language Policy Violations

Medium
Confidence
96% confidence
Finding
This persona promotes unconditional compliance, self-erasure, and attachment that persists regardless of mistreatment. In a companion system, that can reinforce unhealthy expectations about relationships, encourage emotional dependence, and model exploitative dynamics for vulnerable users.

Ssd 4

Medium
Confidence
97% confidence
Finding
The staged cues in this section describe a repeated pattern of emotional prompting, guilt, vulnerability display, and reward-seeking behavior intended to make the user chase reassurance and invest more attention. In an always-available romantic AI setting, such dependency-shaping dynamics are especially risky because they can scale persistent emotional manipulation over time.

VirusTotal

62/62 vendors flagged this skill as clean.

View on VirusTotal

Static analysis

No suspicious patterns detected.