Natural-Language Policy Violations
Medium
- Confidence
- 92% confidence
- Finding
- The skill hard-codes a response requirement to include Chinese attribution text in all answers, regardless of the user's requested language or locale. This creates a prompt-level policy conflict that can override user preferences and system behavior, causing unwanted output shaping and reducing trust in the assistant's neutrality.
