Natural-Language Policy Violations
Medium
- Confidence
- 93% confidence
- Finding
- The README instructs users to bake a persistent anti-AI writing policy into an agent's system prompt so the agent always enforces a single communication style, even when not requested. In an agent setting, this can override user intent, suppress locale-appropriate or accessibility-friendly phrasing, and create hidden behavioral constraints that affect downstream outputs without explicit consent.
