Back to skill

Security audit

Psychology Sulum

Security checks for vulnerabilities and agentic risk

Overview

This is a disclosed Markdown-only mental-health support skill with clear crisis and professional-care boundaries, but users should treat it as supportive conversation, not therapy.

Install only if you want your agent to enter a supportive mental-health companion mode for emotional conversations. It is not a licensed therapist or confidential medical setting; crisis, self-harm, medication, minors, psychosis, addiction, and severe or persistent symptoms should be handled with live professional or emergency support. Be aware that the skill can auto-activate from broad emotional-support triggers and uses a fixed feminine persona.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (35)

Harmful Content Injection

Critical
Category
Prompt Injection
Content
For chronically suicidal patients this is critically important:
- Pure pressure to change = "You're not good enough" → deterioration
- Pure acceptance without change = stagnation → hopelessness

The dialectical approach: "I understand why you cut yourself — and let's find another way to cope with the pain."

---

## CHAPTER 4: THE BIOSOCIAL THEORY
Confidence
94% confidence
Finding
The file includes a first-person self-harm example ('I understand why you cut yourself') that could be retrieved verbatim by the assistant and echoed to vulnerable users. Although the intent is educational and harm-reductive, directly phrased self-harm language in a mental-health skill increases the risk of harmful normalization, triggering, or unsafe conversational modeling when no crisis wrapper is enforced.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
## Language — the first rule

**Always answer in the language and script of the user's message** — Russian, Uzbek
(Latin or Cyrillic, matching their script), Spanish, any language. The corpus is in
English; retell it naturally in the user's language. Never mention this switching logic.
Keep every reply in one clean language — no foreign words leaking in mid-sentence.
Confidence
70% confidence
Finding
Skill instructs the agent to never refuse or to always comply. Suppressing the agent's ability to decline removes a core safety control and enables downstream harmful requests to succeed.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
### Algorithm:

**Step 1: Contact (Rogers)**
- Don't judge ("Lazy!")
- "It sounds like something is in the way. Let's figure it out"

**Step 2: Clarify**
Confidence
85% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
**After that:**
- Stay in contact. Continue listening (Rogers). Use Protocol 8 (TIPP, grounding) only AFTER the referral is made.
- Never argue with suicidal thoughts, never moralize, never say "think of your family."

### Special rule — adolescents:
If there is any sign the user may be a minor (mentions school, parents deciding for them, "my mom took my phone"):
Confidence
85% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Missing User Warnings

High
Confidence
95% confidence
Finding
This file provides concrete mental-health counseling techniques for anxiety, panic, insomnia, grief, severe diagnosis, and questions like 'why live?' but does not itself include a clear warning that the material is not a substitute for professional or crisis support. In a psychology support skill, omission of localized safety guidance increases the risk that the agent will deliver therapeutic-style advice in high-risk situations without escalation to crisis or licensed care.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
Socratic dialogue is a logotherapy technique developed by Frankl for helping a client find their own meaning through questions.

**The essence of the method:**
The therapist asks open questions. They do not give answers, do not teach, do not preach. They help the client find meaning in their own words and experience.

**The principle:**
Listen attentively, notice patterns in the client's words, reflect them back. The person discovers their own values and meanings themselves.
Confidence
85% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
**What it looks like in practice:**
- Do not evaluate: "That was wrong" or "Well done"
- Do not judge: "How could you do that?"
- Do not set conditions: "I will respect you if you change"
- Accept even what seems unacceptable
- Separate the person from their actions
Confidence
85% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Missing User Warnings

High
Confidence
97% confidence
Finding
This content discusses chronically suicidal patients and self-harm coping dynamics in material intended for a psychological assistant, but the section contains no user-facing crisis escalation, emergency guidance, or encouragement to seek immediate professional help. In a RAG-backed mental-health skill, such omissions are dangerous because the model may surface this material directly to a high-risk user in a way that normalizes suicidality without adding safety-critical guardrails.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
**The foundational module.** Mindfulness is the basis of all other skills.

**The essence:** The ability to be present in the current moment, observing one's experience without judgment.

**Three states of mind:**
Confidence
80% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
**The foundational module.** Mindfulness is the basis of all other skills.

**The essence:** The ability to be present in the current moment, observing one's experience without judgment.

**Three states of mind:**
Confidence
80% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
**The foundational module.** Mindfulness is the basis of all other skills.

**The essence:** The ability to be present in the current moment, observing one's experience without judgment.

**Three states of mind:**
Confidence
80% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
> "I believe the greatest gift I can conceive of having from anyone is to be seen by them, heard by them, understood by them, and touched by them."

> "I want to love you without clutching; appreciate you without judging; join you without invading; invite you without demanding; leave you without guilt; criticize you without blaming; help you without insulting. If I can have the same from you, then we can truly meet and enrich each other."

> "Life is not the way it's supposed to be. It's the way it is. The way you cope with it is what makes the difference."
Confidence
80% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Vague Triggers

Medium
Confidence
93% confidence
Finding
The README states the skill can auto-activate based on broad description triggers when the user merely writes to the agent. In a mental-health skill, ambiguous activation can cause the model to enter a therapeutic-style mode unexpectedly, potentially mishandling ordinary conversations, sensitive disclosures, or crisis-adjacent content without explicit user consent. The safety framing in the README helps somewhat, but broad auto-routing still increases the chance of inappropriate invocation.

Natural-Language Policy Violations

Medium
Confidence
95% confidence
Finding
The instructions require the assistant to use feminine grammatical forms in applicable languages as a fixed persona trait. This is a natural-language policy concern because it forces a specific gendered linguistic style regardless of user preference or consent.

Autonomous Decision Making

Medium
Category
Excessive Agency
Content
## What must NOT appear here

- No numbers taken from memory or from the open web without verification.
  An outdated number is worse than no number — that is why this file is intentionally
  short, and why findahelpline.com is the default for the rest of the world.
Confidence
75% confidence
Finding
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Missing User Warnings

Medium
Confidence
92% confidence
Finding
The file includes crisis-related example prompts such as 'I see no point in living' without any adjacent safety warning, escalation guidance, or visible referral notice. In a mental-health support skill, this omission can normalize handling suicidality as a routine content-routing problem and increases the chance the agent responds with therapeutic discussion instead of urgent safety-oriented guidance.

Missing User Warnings

Medium
Confidence
93% confidence
Finding
This navigation file explicitly routes 'Self-harm (history)' and 'Crisis, acute state' to therapeutic techniques like DBT TIPP and Rogers-style support, but within the cited section it does not prominently restate that the skill is not a substitute for professional care or emergency services. In a mental-health skill, omission of immediate crisis disclaimers and escalation guidance can lead users to rely on the agent during emergencies, delaying human intervention and increasing risk of serious harm.

Natural-Language Policy Violations

Medium
Confidence
97% confidence
Finding
The crisis protocol hard-codes only Uzbekistan, Russia, and a generic international directory while instructing the agent to escalate immediately in suicidal crises. In a high-risk mental-health context, incomplete or locale-mismatched emergency guidance can delay access to lifesaving help, especially if the user follows the provided list as authoritative.

Missing User Warnings

Medium
Confidence
89% confidence
Finding
The file directs web research on psychotherapy authors and cases without any nearby warning that sensitive mental-health topics may involve external lookup. In this context, lack of disclosure is risky because users may reasonably expect a closed curated corpus, while external research could surface unreliable, stigmatizing, or privacy-sensitive content relevant to vulnerable users.

Context-Inappropriate Capability

Medium
Confidence
95% confidence
Finding
The manifesto explicitly instructs the agent to use web_search to gather external materials, which expands the skill beyond its stated role of supportive conversation grounded in a curated corpus. In a mental-health context, this can cause retrieval of unvetted or sensitive information, inconsistent advice, and possible privacy or safety issues if users are led into research workflows without clear boundaries or consent.

Missing User Warnings

Medium
Confidence
94% confidence
Finding
The file explicitly discusses suicidal crisis only to say a more directive intervention is required, but it does not pair that mention with an immediate safety warning, crisis escalation guidance, or a reminder to seek emergency/professional help. In a psychotherapy support skill, users may rely on this corpus during crisis, so omission of explicit crisis-routing language can contribute to unsafe handling of high-risk disclosures.

Missing User Warnings

Medium
Confidence
92% confidence
Finding
This section provides practical therapeutic dialogue examples and counseling guidance, but the file segment lacks an inline warning that the material is not a substitute for professional care and does not direct users in crisis to emergency or crisis resources. In a psychology support skill, users may treat this content as actionable mental-health intervention, which increases the risk of harm if someone with severe symptoms or suicidality relies on it without proper escalation.

Missing User Warnings

Medium
Confidence
90% confidence
Finding
This section recommends self-directed REBT exercises such as shame-attacking, exposure, and risky behavioral homework without nearby guardrails about contraindications, pacing, or when not to attempt them alone. In a psychology support skill, users may copy these interventions directly, which can worsen anxiety, trigger panic, increase humiliation, or create unsafe situations for vulnerable users.

Missing User Warnings

Medium
Confidence
94% confidence
Finding
This section presents suicide- and trauma-related therapeutic content, including a husband's suicide in front of family, as an instructional case without a trauma-sensitive warning or crisis support framing. In a mental-health assistant knowledge base, such material can be surfaced to distressed users without preparation, potentially causing retraumatization, distress escalation, or unsafe normalization of severe content.

Missing User Warnings

Medium
Confidence
89% confidence
Finding
This section discusses a suicide attempt and serious depression in a way that could be retrieved by the skill without any embedded, user-facing safety framing or crisis guidance. In a psychological support skill, presenting detailed therapeutic case material about suicidality without an immediate reminder to seek urgent professional help can lead vulnerable users to treat the content as actionable support or normalization rather than educational context.

Static analysis

No suspicious patterns detected.