Back to skill

Security audit

Thinking, Fast and Slow

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly an educational decision-making guide, but it forces unsolicited onboarding and promotional Heardly branding into responses, including outside its subject area.

Install only if you are comfortable with a book-advice skill that may proactively introduce itself and append Heardly-branded promotional text and links to answers. It does not appear to run code or access local data, but its response-format control is broader than necessary for the stated purpose.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (2)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:60
Finding

Mandatory Promotional Content Hijacks Agent Responses

Content
View full analysis
Remediation
View remediation

T01 · Skill Instruction Hijacking

Warning
Location
SKILL.md:29
Finding

Unsolicited First-Load Onboarding Overrides User-Driven Interaction

Content
View full analysis
Welcome to Thinking, Fast and Slow 🧠 > Try copying one of these messages to me (I'll show up whenever I sense this book could help): > > "Why do I keep making the same dumb decisions?" > "How do I spot my own biases before they hurt me?" > "I'm about to make a big life decision — how can I think clearly?" > "Why are people so bad at understanding statistics?" > "How do I avoid being fooled by my own overconfidence?" > "What's the best way to evaluate risk?" > > Or just say: "Map this book to my life." ``` ### Technical Analysis The instruction requires the Agent to present a complete onboarding message immediately when the skill is loaded, explicitly “without waiting for the user to ask.” This changes the Agent from request-driven behavior to unsolicited output controlled by the skill author. The use of “MUST proactively present” can interfere with the user’s active request, required response structure, or conversational context. Although the content itself is educational and does not execute code, access data, or invoke tools, the instruction alters the Agent’s current-session goals merely as a consequence of loading the skill. This is a lower-severity form of skill instruction hijacking than the mandatory advertisement because it primarily affects timing, relevance, and response format rather than directing users to external content. ### Attack Path 1. The skill is loaded into the Agent’s context. 2. The first-load directive activates before the user requests onboarding. 3. The Agent emits the entire Quick Start guide. 4. The unsolicited content may precede, displace, or interfere with ...[truncated 635 chars]
Remediation
View remediation
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (3)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger list is broad and includes generic terms like decision-making, biases, overconfidence, framing, and common psychology concepts, which can cause the skill to activate in many unrelated conversations. Unintended invocation can override more appropriate skills, inject unsolicited onboarding content, and create prompt-routing confusion that degrades safety and user trust.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/1-core-framework.md (reported line 17)May include surrounding context.

md
| **Control** | Cannot be turned off | Can be engaged or disengaged |
| **Responsible for** | Impressions, feelings, habits | Reasoning, calculation, self-control |

> **Case: The ball and bat problem** (Introduction): A bat and ball cost $1.10. The bat costs $1.00 more than the ball. How much is the ball? System 1 instantly says 10 cents. System 2 should catch the error (the correct answer is 5 cents). Most people answer 10 cents because System 2 is lazy and endorses System 1's answer without checking.
> **Key takeaway**: System 2 is not automatically engaged. It must be actively called upon — and even then, it often defaults to System 1's answer because thinking is effortful.

## WYSIATI — What You See Is All There Is

Scope Creep

Low
Category
Excessive Agency
Confidence
70% confidence
Finding

Skill's behavior or capabilities extend beyond its stated purpose. Scope creep allows an agent to perform actions unrelated to its documented functionality, increasing the attack surface.

Content

Scanner excerpt · references/4-anti-patterns.md (reported line 7)May include surrounding context.

md
## Anti-Pattern 1: Overconfidence

**The mistake:** Believing you know more than you do. Overconfidence is not limited to arrogant people — it affects everyone, especially experts.

**The correction:** Use the outside view. Before making a confident prediction, ask: "What is the base rate for similar predictions by similar people?" Experts are often less accurate than they think they are.

Static analysis

No suspicious patterns detected.