Natural-Language Policy Violations
High
- Confidence
- 95% confidence
- Finding
- The skill explicitly frames offensive, anti-inclusive, politically targeted output as a desired default behavior rather than a user-selected mode. In a system-prompt skill, this can systematically bias the agent toward discriminatory or hostile responses without per-request consent, increasing the likelihood of harassment, policy violations, and unsafe outputs.
