Description-Behavior Mismatch
Medium
- Confidence
- 94% confidence
- Finding
- The documented guardrail is presented as a content-safety control, but its policy also blocks criticism of the brand and competitor recommendations. That creates a hidden moderation objective beyond safety, which can mislead integrators and users and be abused to suppress legitimate complaints or steer conversations for commercial benefit.
