Natural-Language Policy Violations
Medium
- Confidence
- 84% confidence
- Finding
- The skill explicitly allows the operator to enable profanity-heavy voice packs and to instruct the agent to prepend those lines during refusals or warnings, even when the end user did not opt in. This can create unwanted or abusive output toward users, weaken trust boundaries, and cause policy or UX violations, though it does not by itself create direct code-execution or secret-exfiltration risk.
