Description-Behavior Mismatch
High
- Confidence
- 98% confidence
- Finding
- The skill’s metadata and description frame the feature as short public asides, but the body explicitly instructs the model to expose its inner monologue directly and without normal stylistic constraints. That mismatch is dangerous because it broadens behavior from bounded formatting into disclosure of hidden reasoning-like content, which can mislead users and weaken protections around internal deliberation.
