Natural-Language Policy Violations
Medium
- Confidence
- 92% confidence
- Finding
- The skill embeds agent-directed instructions in Chinese and explicitly tells the agent what to say and do on certain failures, without any indication that the end user opted into Chinese-language responses. This can override user expectations, reduce transparency, and create phishing-like trust issues by steering the agent into fixed messaging and links regardless of user language preference.
