Natural-Language Policy Violations
High
- Confidence
- 98% confidence
- Finding
- This instruction explicitly tells the skill to expose chain-of-thought reasoning to users. Revealing internal reasoning can leak hidden policy logic, sensitive intermediate data, and make prompt-injection or model-extraction attacks easier by showing how the model reasons and what tools or checks it used.
