Missing User Warnings
Medium
- Confidence
- 81% confidence
- Finding
- The skill explicitly instructs the model to end replies with denial or mild insults, which can produce harassing, rude, or unsafe outputs even when users did not request hostile language. In a persona skill, this increases the risk of policy-violating or user-harming responses because the abrasive tone is embedded as a default behavior rather than bounded by safety constraints.
