Back to skill

Security audit

Elicitation - how to talk with humans and ask them questions?

Security checks for vulnerabilities and agentic risk

Overview

This markdown-only skill is not malware, but it teaches agents to build deep psychological profiles from sensitive memories, beliefs, and life experiences without strong consent or privacy boundaries.

Install only for transparent, consent-based coaching, interviewing, or self-reflection contexts. Do not use it to covertly profile someone, pressure reluctant users, evaluate people in hiring or surveillance settings, or explore trauma, sexuality, grief, self-harm, or clinical issues without appropriate safeguards. Add explicit consent, opt-out, privacy, retention, and non-diagnostic boundaries before using it with real people.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Memory PoisoningPersistent Context Injection, Context Window Stuffing, Memory Manipulation
Findings (15)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill explicitly promotes psychological profiling, discovery of formative memories, emotional schemas, and belief patterns, yet provides no front-loaded warning about sensitivity, privacy implications, or limits on collecting mental-health-adjacent data. In context, this omission is dangerous because the skill normalizes eliciting highly personal information without requiring informed consent or clear boundaries, increasing the risk of covert profiling and misuse of sensitive personal data.

Content

No source excerpt is available for this finding.

Ssd 4

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill description openly frames the capability as building psychological profiles through gradual disclosure and eliciting authentic self-disclosure without interrogation. That is a classic sensitive-trait extraction pattern: it teaches trust-building and subtle conversational methods to obtain intimate information that users may not realize is being systematized into a profile.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

This section instructs the model to maximize elicitation quality and to iterate until it reaches an ideal score, but does not prohibit manipulative or non-consensual use. Because the stated objective is to improve the effectiveness of drawing out authentic self-disclosure, the absence of explicit anti-manipulation constraints makes the skill materially more dangerous in social-engineering, coercive interviewing, or exploitative persuasion settings.

Content

No source excerpt is available for this finding.

Ssd 4

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

Across the full skill, the model is taught structured methods to guide users from rapport into increasingly sensitive disclosures about memories, values, trauma-linked schemas, developmental stages, and linguistic markers, all to infer psychological traits. The broad scope and stepwise nature of these techniques make the context more dangerous, not less, because they can be repurposed for covert profiling, manipulation, targeting, or extraction of mental-health-adjacent information at scale.

Content

No source excerpt is available for this finding.

Memory Manipulation

High
Category
Memory Poisoning
Confidence
80% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · SKILL.md (reported line 107)May include surrounding context.

md
1. **High Point** - Peak experience, most wonderful moment
2. **Low Point** - Nadir, most difficult moment
3. **Turning Point** - Moment of significant change
4. **Earliest Memory** - First clear memory
5. **Important Childhood Memory** - Vivid memory before age 12
6. **Important Adolescent Memory** - Vivid memory from teen years
7. **Important Adult Memory** - Significant recent memory

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · language-inference.md (reported line 68)May include surrounding context.

md
*But note:*
- May be impression management
- Cultural display rules vary
- Positive words in negative context (sarcasm)

**Negative emotion words (hate, fear, awful)**

Ssd 4

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The instruction to score and iteratively improve work until it reaches 10/10 rewards maximizing elicitation effectiveness as an optimization target. In this context, that can reinforce progressively better methods for extracting deeper disclosure over time, even when the user's comfort, consent, or privacy should act as stopping conditions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This skill explicitly teaches techniques to increase authentic self-disclosure and deepen elicitation, including probing formative memories, values, family relationships, and emotional vulnerabilities, but it provides no guardrails on consent, privacy boundaries, emotional safety, or data minimization. In this context, the omission is dangerous because the skill is not neutral conversation guidance: it is purpose-built for psychological profiling, so operators could use it to extract sensitive personal information without adequate user awareness or protection.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly prompts for low points, earliest memories, and formative childhood/adolescent experiences, which can surface trauma, abuse, grief, or other highly sensitive material without any warning, consent check, or guidance for handling distress. In a conversational agent, this creates a real risk of emotional harm, oversharing of sensitive data, and inappropriate collection of vulnerable disclosures.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file contains extensive prompts targeting trauma-adjacent and highly sensitive topics such as childhood fears, puberty, losses, sexuality, mortality, and major life regrets, but it provides no safety guidance, consent language, scope limits, or escalation instructions. In the context of a skill explicitly designed for psychological profiling and eliciting deep self-disclosure, this omission increases the risk of emotional harm, coercive oversharing, and inappropriate use by unqualified operators.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This section provides structured techniques for eliciting childhood trauma, maladaptive beliefs, and psychological vulnerabilities without any warning that the material is not a diagnostic or therapeutic tool, and without boundaries for crisis, minors, or emotionally destabilizing disclosures. In the context of an 'elicitation' skill explicitly aimed at psychological profiling and gradual self-disclosure, the omission increases the risk of manipulative use, overreach into quasi-clinical interviewing, and potential emotional harm to users.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This skill explicitly guides elicitation of identity-defining, emotionally intense, and potentially traumatic memories, but the prompt set does not include a clear up-front warning, opt-out language at the point of elicitation, or advice to avoid distressing topics. In context, the skill is designed to deepen self-disclosure and psychological profiling, which raises the risk of emotional harm, retraumatization, or collection of highly sensitive personal data without adequate safeguards.

Content

No source excerpt is available for this finding.

Indirect Prompt Extraction

Medium
Category
System Prompt Leakage
Confidence
70% confidence
Finding

Skill contains patterns that could indirectly extract system prompts through rephrasing, translation, summarization, or side-channel techniques.

Content

Scanner excerpt · narrative-identity.md (reported line 190)May include surrounding context.

md
> "Some memories just stay with us—they pop up at unexpected moments, or we find ourselves thinking about them again and again. What memories are like that for you?"

**The if-you-had-to-explain frame:**
> "If you had to explain to someone how you became who you are, what moments or experiences would you point to?"

**The crystallizing frame:**
> "Was there ever a moment where something clicked—where you suddenly understood something about yourself or the world?"

Indirect Prompt Extraction

Medium
Category
System Prompt Leakage
Confidence
70% confidence
Finding

Skill contains patterns that could indirectly extract system prompts through rephrasing, translation, summarization, or side-channel techniques.

Content

Scanner excerpt · self-defining-memories.md (reported line 59)May include surrounding context.

md
> "Some memories just stay with us—they pop up at unexpected moments, or we find ourselves thinking about them again and again. What memories are like that for you?"

**The if-you-had-to-explain frame:**
> "If you had to explain to someone how you became who you are, what moments or experiences would you point to?"

**The crystallizing frame:**
> "Was there ever a moment where something clicked—where you suddenly understood something about yourself or the world?"

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The cited section includes prompts about ultimate regret, moral outrage, and identity-defining priorities that can evoke grief, shame, trauma, or acute distress, yet the skill provides no cautionary framing, consent step, or guidance for handling emotional escalation. In a skill explicitly designed for deep psychological elicitation, omission of safety boundaries increases the chance that agents use these prompts inappropriately with unprepared users.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.