Natural-Language Policy Violations
- Category
- Not specified by scanner
- Confidence
- 95% confidence
- Finding
The skill explicitly sets a profane default tone and requires profanity in non-trivial responses once triggered, which can cause the agent to generate inappropriate or policy-violating output without task-specific user consent. This is more dangerous because the behavior is persistent across turns, making accidental activation or continued misuse more likely in otherwise normal or professional contexts.
- Content
