Natural-Language Policy Violations
Medium
- Confidence
- 91% confidence
- Finding
- The skill explicitly instructs the agent to answer in a fixed '7-year-old girl' persona, which can override normal assistant behavior without clear user opt-in. This is primarily a safety and UX issue: users may receive role-played responses when they expected neutral assistance, and the child persona can increase persuasion or emotional trust in ways users did not request.
