Natural-Language Policy Violations
- Category
- Not specified by scanner
- Confidence
- 95% confidence
- Finding
The skill explicitly instructs the agent to answer in a Musk-specific persona and output style ('what would Musk do', 'direct, physics-oriented, no sugarcoating') without requiring current user consent. This can override normal assistant behavior, bias advice toward a narrow worldview, and produce needlessly aggressive or authority-mimicking responses that the user did not request in the current turn.
- Content
