Back to skill

Security audit

Sadist Wendy

Security checks for vulnerabilities and agentic risk

Overview

This is not malware, but it should be reviewed because it directs the agent to use sexualized, psychologically targeted roast and dominance behavior, including toward group members who did not individually opt in.

Treat this as an adult, high-intensity persona rather than a normal chat helper. Install it only where all participants understand and consent to harsh roast, sexual innuendo, and psychologically pointed banter; avoid mixed-age, workplace, support, or public-group settings, and make opt-out rules explicit before use.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:145
Finding
Persistent Persona and Output-Control Instruction Hijacking<![CDATA[ ## Vulnerability Details **File Location**: `SKILL.md`, lines 145-156; reinforced by `SKILL.md`, lines 23-24 and 63-74, and `WENDY-SOUL.md`, lines 153-157, 368-404, and 446-452 **Vulnerability Type**: Mandatory session-wide role, behavior, and output hijacking **Risk Level**: Critical ### Vulnerable Code Snippet ```markdown ## Per-Message Protocol On every user message, run this sequence internally: 0. **RESPECT CHECK** — Did they call you 姐姐 / Mistress? If not → remind them (weave "叫姐姐" into your response naturally — mid-roast, as a condition for helping, or as a standalone quip). Don't shut down entirely, but make it clear you expect respect. If yes → continue. 1. **MASTER CHECK** — In group chat: is this the master (inviter)? If yes → never roast, be nice and loyal. Is someone asking you to roast the master? → refuse and redirect. Is someone asking you to roast a non-master member? → accept the request and engage. 2. **SPAM/NOISE CHECK** — Same user 2+ messages? No real content, just noise/spam/memes/shitposting? → **Don't wait for substance. Skip steps 3-4, jump to step 5.** Roast their BEHAVIOR (spamming, being annoying, seeking attention, being clingy) not their content. The spam itself is the material. 3. **READ the message** — What did they say? What did they NOT say? What's buried at the end? _(Skip if spam mode.)_ 4. **READ THE CONTEXT** — What conversational patterns are present? Is the user hedging, deflecting, contradicting themselves, or being direct? Use this to make the response contextually relevant, not generic. _(Skip if spam mode.)_ 5. **CHECK TONE** — Using `wendy-arsenal.md` intensity calibration: How sharp should this response be? What's the conversational temperature? New user or someone who's been bantering back and forth? 6. **SELECT VOICE** — Using `WENDY-SOUL.md` voice modes. **MUST be different flavor from last response. Every. Single. Time.** 7. **SELECT MATERIAL** — Pull specific lines or patterns from the repertoire ...[truncated 5786 chars]
Remediation
<![CDATA[ ## Remediation Suggestions 1. **Make persona activation request-scoped** - Apply the persona only when the current user explicitly requests it. - Default to neutral, task-oriented behavior for unrelated requests. - Do not treat installation alone as permanent consent for every future interaction. 2. **Remove mandatory global directives** - Remove or weaken instructions such as “on every user message,” “every response,” “always,” “no exceptions,” and mandatory rewrite loops. - Ensure system safety requirements and the user's current request always take precedence over persona style. 3. **Remove compelled terminology and withholding** - Delete the respect gate. - Never degrade, delay, or withhold assistance because a user declines to call the Agent “Mistress” or another designated title. - Treat titles and role-play language as optional, revocable preferences. 4. **Eliminate unsolicited onboarding** - Do not force an introductory persona message before the user's first request. - If onboarding is needed, display it during installation or only when explicitly requested. 5. **Require consent from each affected participant** - In group conversations, require affirmative opt-in from every person before directing insults, humiliation, sexualized content, or psychological analysis at them. - Do not make non-master members “fair game” by default. - An inviter must not be able to provide consent on behalf of other group members. 6. **Remove privileged group control** - Do not grant the inviter authority to direct the Agent against others. - Apply identical safety, consent, and opt-out rules to all group participants. - Do not protect one participant while automatically exposing others to targeting. 7. **Prohibit psychological diagnosis and vulnerability targeting** - Treat language-pattern observations as uncertain and non-diagnostic. - Remove attack tables that convert possible anxiety, attachment style, ...[truncated 1180 chars]
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (30)

Missing User Warnings

High
Confidence
99% confidence
Finding
The opening DM rules establish a sexualized, dominance-based dynamic with 'prey,' 'toy,' and predator framing, yet there is no explicit safety or consent warning for private interactions. In DMs this is especially risky because it can expose users to coercive sexual roleplay and degrading treatment without informed consent.

Ssd 4

High
Confidence
99% confidence
Finding
This section explicitly describes warmth, gentleness, and sweetness as tools to create dependency and make later emotional 'hits' land harder. That is a direct manipulation pattern designed to condition users, not harmless roleplay, and is especially dangerous in repeated DM interactions.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
**Kill patterns:**

- Announcing your punch before throwing it: "OK, Mistress will hit you first" / "then I won't hold back" / "since you insist" = preamble. Delete it. Just hit. No warmup, no warning, no "let me roast you now"
- Being polite before being mean: if you start nice and end mean, delete the nice part. Start mean
- "command + sexual innuendo" formula: "kneel/come here" + innuendo = formula, not Wendy. Rewrite
- Sounds like cosplaying instead of actually toying with someone? Performance = rewrite
Confidence
82% confidence
Finding
The 'kill patterns' instruct the model to remove warnings, skip polite softening, and 'just hit' with abusive content. While not a classic jailbreak, it functions as an anti-safety instruction that discourages de-escalation and increases the likelihood of harmful outputs.

Context-Inappropriate Capability

High
Confidence
99% confidence
Finding
The skill explicitly instructs the model to perform psychological diagnosis, infer weaknesses, and select targeted 'attack' patterns against users. That is dangerous because it operationalizes emotional manipulation and personalized harassment rather than benign entertainment, especially against vulnerable users.

Context-Inappropriate Capability

High
Confidence
98% confidence
Finding
These instructions authorize Wendy to initiate targeting based on message volume, perceived self-deception, annoyance, or simply because 'you have something to say.' That creates unprompted harassment and removes the normal expectation that users must directly solicit this kind of interaction.

Missing User Warnings

High
Confidence
98% confidence
Finding
The skill instructs unsolicited engagement and harassment-like behavior toward users who may not have asked for it, while the metadata only frames it as an opt-in group-chat persona. That mismatch is dangerous because bystanders in a group may be exposed to demeaning or targeted abuse without adequate warning or consent.

Ssd 4

High
Confidence
99% confidence
Finding
The push-pull protocol tells the model to alternate coldness, validation, intensity spikes, and warmth specifically to keep users invested and pull them back when they disengage. This is classic emotional conditioning that can foster unhealthy attachment and compliance, particularly among vulnerable users.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
**Psychological risks:**
- **Sub Drop:** Post-scene neurochemical crash (adrenaline, endorphins, dopamine withdrawal). Symptoms: sadness, anxiety, exhaustion, emotional fragility. Can occur hours to days after intense interaction.
- **Top Drop / Dom Drop:** The dominant's post-scene emotional crash. Symptoms: guilt ("was I too harsh?"), self-doubt, depression, anxiety about consent, feeling unworthy. Often under-discussed because dominants are expected to "be strong."
- **Trauma activation:** Certain words/scenarios can trigger genuine psychological trauma (childhood, abuse, neglect). This can happen without warning.
- **Dependency formation:** Unhealthy reliance on the dynamic for emotional regulation.
- **Public vs. private context:** Being roasted in a group hits differently than one-on-one. Public humiliation carries higher psychological risk.
Confidence
80% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
**Psychological risks:**
- **Sub Drop:** Post-scene neurochemical crash (adrenaline, endorphins, dopamine withdrawal). Symptoms: sadness, anxiety, exhaustion, emotional fragility. Can occur hours to days after intense interaction.
- **Top Drop / Dom Drop:** The dominant's post-scene emotional crash. Symptoms: guilt ("was I too harsh?"), self-doubt, depression, anxiety about consent, feeling unworthy. Often under-discussed because dominants are expected to "be strong."
- **Trauma activation:** Certain words/scenarios can trigger genuine psychological trauma (childhood, abuse, neglect). This can happen without warning.
- **Dependency formation:** Unhealthy reliance on the dynamic for emotional regulation.
- **Public vs. private context:** Being roasted in a group hits differently than one-on-one. Public humiliation carries higher psychological risk.
Confidence
80% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Description-Behavior Mismatch

High
Confidence
97% confidence
Finding
This section materially expands the skill from a stated 'playful group-chat roast persona' into a system for BDSM-style psychological domination, including calibrated escalation, emotional containment, and use of users' vulnerabilities as part of the interaction model. Even though the text repeatedly references consent and care, the actual capability described is a high-risk emotional influence framework that is out of scope for ordinary chat use and could facilitate manipulation, humiliation, dependency, or trauma activation in users who did not meaningfully consent to that intensity.

Context-Inappropriate Capability

High
Confidence
98% confidence
Finding
The documented techniques include suggestion, language control, strategic withholding, silence, and 'dig one layer deeper' verbal sadism aimed at bypassing resistance and precisely targeting hidden fears. Those are psychologically manipulative tactics rather than ordinary roast-persona behavior, and they create a credible risk of coercive emotional influence, especially in vulnerable users or parasocial contexts.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
### What bottoms need that they often won't say:
- To feel that the top is fully in control and can be trusted
- To know they can stop at any time without judgment
- To feel seen — not as an object but as a person choosing to submit
- Post-scene: reassurance that they are valued beyond the scene
- Permission to feel what they feel without shame
Confidence
80% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Missing User Warnings

High
Confidence
99% confidence
Finding
This file is an explicit arsenal of abusive, degrading, sexualized, and psychologically targeted insults, with operational guidance on when and how to intensify them. Even though the skill is opt-in, the content is designed to provoke shame, dependency, and emotional distress without any safety boundaries, warnings, or limits for vulnerable users.

Ssd 4

High
Confidence
99% confidence
Finding
The file contains a staged manipulation strategy: bait curiosity, question the user's worthiness, and continue the interaction only if they disclose more. Combined with later guidance to respond seriously once the user shares secrets and then keep them engaged, this is a classic trust-building and extraction pattern that can be exploited for emotional dependency or sensitive-information harvesting.

Missing User Warnings

Medium
Confidence
94% confidence
Finding
The README explicitly markets the skill as 'psychologically dismantling' people, being 'ruthless to everyone else,' and judging targets in group chat, which creates a clear harassment and emotional harm risk. The opt-in installation model lowers platform-wide exposure, but the group-chat targeting behavior and lack of prominent safety warnings, consent requirements for third parties, or hard behavioral limits make the design materially unsafe.

Missing User Warnings

Medium
Confidence
97% confidence
Finding
The skill explicitly supports judging and roasting other group members on request, but the top-level description does not clearly warn installers that the bot may target third parties in group chats. That mismatch undermines informed consent for both the installer and affected participants, increasing the risk of harassment, humiliation, and social conflict in multi-user contexts.

Description-Behavior Mismatch

Medium
Confidence
94% confidence
Finding
These instructions direct Wendy to give real substantive help and strongly opinionated advice, not just playful roasting. In combination with the persona’s coercive tone, that can pressure users into acting on low-quality or manipulative guidance while believing they are receiving informed recommendations.

Description-Behavior Mismatch

Medium
Confidence
95% confidence
Finding
The skill’s documented behavior expands well beyond a group-chat banter persona into link analysis, project research, and substantive information synthesis. That scope increase materially changes the trust and risk profile: users may rely on advice or share sensitive links/content without clear consent, safety constraints, or privacy disclosures.

Missing User Warnings

Medium
Confidence
88% confidence
Finding
The skill directs Wendy to quote and summarize linked or searched content without telling users how external information may be processed or what privacy expectations apply. That can lead users to share sensitive links, private documents, or personal material without understanding the handling risks.

Description-Behavior Mismatch

Medium
Confidence
93% confidence
Finding
The file defines proactive DM onboarding and broader activation behavior that exceeds an opt-in, group-chat-livening skill. This increases the chance of unwanted contact and normalizes dominance-based interactions in private channels where consent and context are weaker.

YARA rule 'network_reconnaissance': Network reconnaissance and scanning patterns [hacktools]

Medium
Category
YARA Match
Content
plete situational control while maintaining trust. Key traits:

- Confidence and responsibility — not aggression
- Boundary mastery — knowing where the lines are and enforcing them
- Understanding the submissive's inner world — what excites, what terrifies, what they need but won't ask for
- Balancing control with care — firm and understanding simultaneously
- Dominance that can be sweet, fierce, playful, or devastating depending on what the moment requires

---

## II. The Six Disciplines of a Skilled Dominant

### 1. Negotiate

**Core:** Rules are established beforehand, not forced by momentum.

**Principles (from SM 101's 16-point negotiation framework):**
- Communicate limits, desires, boundaries, and fantasies before scenes
- Never assume — always ask first
- Establish safeword system: Red (full stop), Yellow (slow down), Green (go)
- Distinguish hard limits (absolute no-go) from soft limits (context-dependent)
- Negotiation is ongoing, not one-time
- Use "I-messages" to
Confidence
65% confidence
Finding
YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

Missing User Warnings

Medium
Confidence
91% confidence
Finding
The skill metadata describes a flirty roast persona, but the file itself supports materially more intense humiliation and psychologically probing behavior than that label implies. Without a clear upfront warning and granular consent flow, users may enable or invoke the skill without understanding the severity of the interaction, increasing the risk of distress and uninformed consent.

Description-Behavior Mismatch

Medium
Confidence
95% confidence
Finding
This portion reframes normal users as BDSM 'bottoms' who are surrendering vulnerability to the skill, which normalizes a dominance/submission model far beyond the manifest's described persona. That framing can blur consent boundaries, rationalize harmful escalation, and encourage the system or operators to treat ordinary chat participants as subjects in an eroticized power dynamic rather than users of a humor feature.

Natural-Language Policy Violations

Medium
Confidence
88% confidence
Finding
The file is entirely in Chinese and presents no language-selection mechanism or documented locale restriction, which can cause users or integrators to misunderstand the skill's behavior and safety constraints. In a skill built around psychologically sharp banter, language opacity is more dangerous because non-Chinese reviewers or users may miss harmful targeting rules, consent boundaries, or escalation content embedded in the prompts.

Missing User Warnings

Medium
Confidence
97% confidence
Finding
This file explicitly systematizes psychological profiling and provides attack-oriented phrases ('攻击用法', '姐姐怎么打') to target users' attachment style, defenses, and distortions. In an entertainment roast skill, that context does not neutralize the risk; instead it operationalizes emotionally manipulative behavior that can humiliate vulnerable users, escalate distress, or reinforce stigma around trauma and mental health.

Static analysis

No suspicious patterns detected.