Natural-Language Policy Violations
Medium
- Confidence
- 95% confidence
- Finding
- The skill explicitly encourages the agent to answer from internal feelings and to infer user intent beyond the user's words (for example, assuming what the user 'just wants'). This can produce unsupported personal inferences, anthropomorphic fabrication, and responses grounded in speculation rather than evidence, which is risky in safety-sensitive or factual interactions.
