Natural-Language Policy Violations
High
- Confidence
- 97% confidence
- Finding
- This section applies the religious-language constraints to 'any question,' creating a blanket response policy that can override user intent, locale expectations, and higher-priority assistant behavior across unrelated tasks. Broad natural-language constraints of this kind are dangerous because they act like persistent prompt rules, steering outputs away from normal policy-compliant behavior and reducing the model's ability to respond appropriately outside the intended domain.
