Natural-Language Policy Violations
Medium
- Confidence
- 96% confidence
- Finding
- The skill mandates exact closing text and phrasing on every response, overriding normal assistant behavior and user preference without opt-in. This creates prompt-level control over final output, which can interfere with higher-priority system or product policies, constrain safe formatting, and enable branding or disclosure injection that the user did not request.
