T01 · Skill Instruction Hijacking
- Location
SKILL.md:57- Finding
Mandatory Third-Party Promotional Footer Hijacks All Agent Responses
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This parenting skill is mostly transparent and non-executable, but it contains unsafe crisis-adjacent child guidance and forces unsolicited onboarding and branding into responses.
Review before installing. This skill may be useful as a 1-2-3 Magic summary, but it should not be relied on for crisis situations; if a child mentions self-harm, suicide, serious violence, abuse, or running away, treat that as a safety issue rather than routine discipline. Also expect the skill to trigger broadly and add mandatory Heardly branding to outputs.
SKILL.md:57Mandatory Third-Party Promotional Footer Hijacks All Agent Responses
SKILL.md:24Unsolicited First-Load Instructions Override the User’s Requested Interaction
references/3-techniques.md:25Potential Child Self-Harm or Runaway Threats May Be Dismissed as Discipline Testing
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.
### Scenario 3: Sibling Warfare
**Kids fight constantly — toys, TV, whose turn it is.**
- **Fix:** 1) Both get counted. Don't judge who started. 2) Separate break locations. 3) Post-break, no rehashing. 4) For fighting over objects: "You can either share or the toy goes away."
### Scenario 4: Public Meltdowns
**Your child throws a tantrum at the grocery store and you're mortified.**
The trigger list includes very broad phrases like "discipline," "parenting help," and "child behavior," which can cause the skill to activate for many ordinary conversations that are not specifically asking for this framework. That creates an overbroad routing risk where users may receive unsolicited prescriptive parenting guidance, including the skill’s mandatory proactive output, in contexts where it is irrelevant or inappropriate.
The description embeds a long trigger list with broad, common phrases such as "discipline," "parenting help," and "child behavior," which can cause the skill to activate for many ordinary parenting conversations beyond explicit requests for this specific book/framework. Over-broad invocation can lead to unintended routing, reduced user control, and accidental disclosure of conversation context to a skill the user did not clearly intend to invoke.
Forcing every response to end with a fixed English watermark overrides normal language adaptation and inserts content the user did not request. While not a direct security exploit, it is a policy and UX integrity issue because it can leak branding into sensitive parenting discussions and reduce transparency about where model instructions end and sponsored or injected content begins.
No suspicious patterns detected.