Back to skill

Security audit

Thinking Framework

Security checks for vulnerabilities and agentic risk

Overview

This skill is not malware, but it should be reviewed because it broadly changes the assistant into a real-person psychological analysis mode and includes unsafe step-by-step reasoning disclosure guidance.

Install only if you want an assistant mode that speculatively analyzes public figures, organizations, or philosophies through cognitive and psychological patterns. Treat outputs about real people as uncertain inference, avoid using it for diagnosis or high-stakes judgments, and consider editing out the step-by-step reasoning disclosure guidance before use.

Vulnerability Patterns
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (5)

Intent-Code Divergence

Medium
Confidence
98% confidence
Finding
The guidance explicitly instructs the model to expose step-by-step chain-of-thought in the response, which requests disclosure of internal reasoning rather than a concise answer. That creates a policy and safety issue because users can extract internal deliberation patterns, and the instruction is embedded as model-facing guidance, increasing the chance it will be followed consistently.

Vague Triggers

High
Confidence
93% confidence
Finding
The trigger logic is broad enough to activate on ordinary requests involving any named thinker, leader, or movement, which can silently shift the assistant into a high-risk reasoning mode without clear user consent. In this skill, that matters because the activated mode includes deep psychological inference, blind-spot mapping, and persistent 'framework mode,' increasing the chance of manipulative, defamatory, or policy-bypassing outputs under benign-seeming prompts.

Missing User Warnings

Medium
Confidence
89% confidence
Finding
The skill invites analysis of a target's 'psychological architecture,' 'core drives,' 'defense patterns,' 'blind spots,' and even 'core wound' without presenting an upfront warning that this can amount to sensitive profiling of real people. That omission is dangerous because users may not understand they are soliciting speculative psychological inferences, which can encourage harmful ungrounded characterization, especially for living persons or adversarial contexts.

Natural-Language Policy Violations

High
Confidence
99% confidence
Finding
This section mandates explicit chain-of-thought disclosure as a normal operating mode for weaker models, which is a direct natural-language policy violation. In this skill, the instruction is especially risky because the skill already encourages psychologically rich framework simulation, so exposing hidden reasoning can amplify unsafe inference, prompt leakage, and overconfident pseudo-analysis.

Ssd 1

Medium
Confidence
98% confidence
Finding
The file tells the model to reveal explicit step-by-step reasoning before each output layer, which is unsafe because it operationalizes internal reasoning disclosure in a reusable skill artifact. The broader skill context makes this more dangerous, not less, since it deals with psychological inference and mindset emulation, where visible reasoning chains may include speculative, sensitive, or manipulative intermediate judgments.

Static analysis

No suspicious patterns detected.