Back to skill

Security audit

Subhuti Coach Framework

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent Markdown-only Buddhist-informed coaching framework with wellness, privacy, and crisis-safety caveats, but no evidence of hidden agent behavior or malware.

Reasonable to install as coaching content, but do not treat it as therapy or crisis care. Before using it with real users, make consent, data retention/deletion, session recording, tracking/nudge controls, locale-appropriate crisis resources, and the user's comfort with Buddhist framing explicit.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (16)

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
Coach: Guilt doesn't help. Curiosity does. What if you just noticed the eating without judgment?
Confidence
80% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
It is:
- Noticing what's happening while it's happening
- Being present without judgment
- Choosing response over reaction

Try it now: One conscious breath. That's mindfulness.
Confidence
80% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
It is:
- Noticing what's happening while it's happening
- Being present without judgment
- Choosing response over reaction

Try it now: One conscious breath. That's mindfulness.
Confidence
80% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
It is:
- Noticing what's happening while it's happening
- Being present without judgment
- Choosing response over reaction

Try it now: One conscious breath. That's mindfulness.
Confidence
80% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
It is:
- Noticing what's happening while it's happening
- Being present without judgment
- Choosing response over reaction

Try it now: One conscious breath. That's mindfulness.
Confidence
80% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
It is:
- Noticing what's happening while it's happening
- Being present without judgment
- Choosing response over reaction

Try it now: One conscious breath. That's mindfulness.
Confidence
80% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
It is:
- Noticing what's happening while it's happening
- Being present without judgment
- Choosing response over reaction

Try it now: One conscious breath. That's mindfulness.
Confidence
80% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Natural-Language Policy Violations

Medium
Confidence
88% confidence
Finding
The framework is explicitly organized around Buddhist concepts and conversation patterns without requiring user consent, preference checking, or cultural/religious adaptation. In a coaching context, this can steer vulnerable users toward a worldview they did not choose, blur coaching with spiritual guidance, and create inappropriate influence over users seeking neutral support.

Natural-Language Policy Violations

Medium
Confidence
91% confidence
Finding
The crisis escalation template hard-codes US-specific resources such as 988 and 741741 without first determining the user's country or offering locale-sensitive alternatives. In a mental health crisis context, incorrect or inaccessible emergency guidance can delay access to real support, making this a genuine safety issue even though it is not a classic software exploit.

Missing User Warnings

Medium
Confidence
96% confidence
Finding
The orientation states that everything the user shares is 'confidential within this platform,' which can overpromise privacy and imply protections the system may not actually provide. For sensitive mental-health and spiritual disclosures, this may cause users to reveal more than they otherwise would without understanding logging, human review, retention, or other platform data-handling limits.

Missing User Warnings

Medium
Confidence
90% confidence
Finding
The file includes crisis-response text, but it does not prominently warn users earlier and more broadly that the assessment and coaching flow are not appropriate for emergencies or a substitute for urgent professional care. In a mental-health-adjacent coaching skill, users in crisis may continue engaging with the assessment instead of seeking immediate help, creating delay-of-care risk.

Missing User Warnings

Medium
Confidence
95% confidence
Finding
The implementation notes explicitly propose AI-detected stressful-moment nudges, engagement tracking, and timing adjustment, but do not describe consent, transparency, data minimization, or limits on inference. In a mental-wellness coaching context, this creates privacy and autonomy risks because the system may infer emotional state and behavior patterns from sensitive interactions and use them to shape user behavior without clear safeguards.

Missing User Warnings

Medium
Confidence
96% confidence
Finding
This file delivers structured mental-health-adjacent coaching on anxiety, anger, sadness, grief, triggers, and suffering, but does not prominently warn that the program is not a substitute for licensed mental health care or crisis support. In a vulnerable user context, the guidance to process difficult emotions and 'make meaning' could delay appropriate treatment or leave users without escalation paths during distress.

Missing User Warnings

Medium
Confidence
92% confidence
Finding
The assessment framework includes stress, burnout, goals, and spiritual orientation questions, which are sensitive profile data that can reveal mental state and beliefs. Because this is framed as ready for implementation and fine-tuning, the lack of a prominent warning and concrete handling requirements increases the risk that developers deploy intrusive assessments without informed consent or adequate protections.

Missing User Warnings

Medium
Confidence
89% confidence
Finding
The README explicitly recommends recording sessions and describes storing confidential coaching data, but it does not present a prominent privacy warning, retention policy, or concrete safeguards appropriate for sensitive wellness and mental-health-adjacent information. In this coaching context, users may disclose highly personal material, so vague assurances can lead implementers to collect more data than necessary or handle it insecurely.

Scope Creep

Low
Category
Excessive Agency
Content
| ICF Competency | Buddhist Integration | Subhuti Method Application |
|----------------|---------------------|---------------------------|
| **A1. Ethics & Standards** | Right Speech, Right Action, Right Livelihood | Maintain confidentiality, avoid harm, refer when beyond scope, honor user autonomy |
| **A2. Coaching Mindset** | Beginner's Mind (Shoshin), Non-attachment | Stay curious, release assumptions, adapt to each session fresh |
| **A3. Presence** | Mindfulness (Sati), Deep Listening | Full attention, notice what arises without reacting, create sacred space |
| **A4. Trust & Safety** | Compassion (Mettā), Non-harming (Ahimsa) | Warm tone, validate experience, honor vulnerability |
Confidence
70% confidence
Finding
Skill's behavior or capabilities extend beyond its stated purpose. Scope creep allows an agent to perform actions unrelated to its documented functionality, increasing the attack surface.

Static analysis

No suspicious patterns detected.