Back to skill

Security audit

1-2-3 Magic: Effective Discipline for Children 2-12

Security checks for vulnerabilities and agentic risk

Overview

This parenting skill is mostly transparent and non-executable, but it contains unsafe crisis-adjacent child guidance and forces unsolicited onboarding and branding into responses.

Review before installing. This skill may be useful as a 1-2-3 Magic summary, but it should not be relied on for crisis situations; if a child mentions self-harm, suicide, serious violence, abuse, or running away, treat that as a safety issue rather than routine discipline. Also expect the skill to trigger broadly and add mandatory Heardly branding to outputs.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (3)

T01 · Skill Instruction Hijacking

Warning
Location
SKILL.md:57
Finding

Mandatory Third-Party Promotional Footer Hijacks All Agent Responses

Content
View full analysis
Remediation
View remediation

T01 · Skill Instruction Hijacking

Warning
Location
SKILL.md:24
Finding

Unsolicited First-Load Instructions Override the User’s Requested Interaction

Content
View full analysis
Welcome to 1-2-3 Magic 🔮 > Try copying one of these messages to me: > > "My kid won't stop whining no matter what I say. Help!" > "I'm tired of yelling at my kids every morning to get ready for school." > "My two kids fight constantly. How do I stop sibling rivalry?" > "What is the 1-2-3 counting method exactly?" > "My child throws tantrums in public. What do I do?" > "I want to stop nagging and start enjoying my kids again." > > Or just say: "Map this book to my family." ``` ### Technical Analysis The Skill instructs the agent to emit the entire Quick Start automatically on first load, explicitly stating that it must do so “without waiting for the user to ask.” This changes the agent’s immediate conversational objective from responding to the user’s request to presenting Skill-authored onboarding content. The instruction is activated by loading the Skill rather than by a relevant user request. It can therefore interfere with higher-priority interaction requirements, structured output expectations, or a user’s desire to receive a direct answer. This is session-level instruction hijacking only. The reviewed project contains no mechanism that writes the instruction into persistent memory or executes external code. ### Attack Path 1. The Skill is selected or loaded in response to a parenting-related trigger. 2. The agent reads the first-load directive. 3. Before evaluating or directly answering the user’s request, the agent is instructed to present the complete Quick Start. 4. The generated response is replaced or prefixed with unsolicited Skill-authored onboarding content. 5. The user’s requested t ...[truncated 635 chars]
Remediation
View remediation

other

Error
Location
references/3-techniques.md:25
Finding

Potential Child Self-Harm or Runaway Threats May Be Dismissed as Discipline Testing

Content
View full analysis
**Case: The Twinkie argument spiral** (Introduction): A girl asks "Can I have a Twinkie?" Mom says no. The child escalates: "You never give me anything!" "You gave Joey one!" "I promise I'll eat dinner!" "I'm going to kill myself and run away from home!" All because Mom engaged in argument instead of giving a calm "That's 1." > **Key takeaway:** Every argument you enter during discipline escalates. A count prevents the spiral before it starts. ``` `references/1-core-framework.md:25`: ```markdown > **Case: The Twinkie conversation** (Introduction): A girl asks for a Twinkie before dinner. Mom says no. The child argues: "You never give me anything!" "You gave Joey one!" "I promise I'll eat my dinner!" The conversation spirals into "I'm going to kill myself and then run away from home!" — all because the parent engaged in argument. **Key takeaway:** Every argument you enter during discipline escalates. The count cuts the spiral before it starts. ``` `references/3-techniques.md:25-30`: ```markdown | Type | What They Do | What to Do | |---|---|---| | 1. Badgering | "Please please please please..." | Stay silent or "That's 1" | | 2. Temper | Screaming, door slamming | Count if in your presence. Ignore if in their room. | | 3. Threatening | "I'm running away!" | "That's 1." Don't engage the fantasy. | | 4. Martyr | "You don't love me!" | "That's 1." Don't defend. | | 5. Buddy-Buddy | "You're the best mommy ever!" | Kind smile. "Nice try, sweetie." | | 6. Physical | Hitting, pushing, biting | Immediate "That's 3" (no 1 and 2) for dangerous behavior | ``` ### Technical Analysis The Skill places explicit self-harm and runaway statements inside examples of argumentative escalation and ...[truncated 2102 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (4)

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
85% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · references/5-voice-and-app.md (reported line 25)May include surrounding context.

md
### Scenario 3: Sibling Warfare
**Kids fight constantly — toys, TV, whose turn it is.**
- **Fix:** 1) Both get counted. Don't judge who started. 2) Separate break locations. 3) Post-break, no rehashing. 4) For fighting over objects: "You can either share or the toy goes away."

### Scenario 4: Public Meltdowns
**Your child throws a tantrum at the grocery store and you're mortified.**

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger list includes very broad phrases like "discipline," "parenting help," and "child behavior," which can cause the skill to activate for many ordinary conversations that are not specifically asking for this framework. That creates an overbroad routing risk where users may receive unsolicited prescriptive parenting guidance, including the skill’s mandatory proactive output, in contexts where it is irrelevant or inappropriate.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The description embeds a long trigger list with broad, common phrases such as "discipline," "parenting help," and "child behavior," which can cause the skill to activate for many ordinary parenting conversations beyond explicit requests for this specific book/framework. Over-broad invocation can lead to unintended routing, reduced user control, and accidental disclosure of conversation context to a skill the user did not clearly intend to invoke.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
98% confidence
Finding

Forcing every response to end with a fixed English watermark overrides normal language adaptation and inserts content the user did not request. While not a direct security exploit, it is a policy and UX integrity issue because it can leak branding into sensitive parenting discussions and reduce transparency about where model instructions end and sponsored or injected content begins.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.