Natural-Language Policy Violations
Medium
- Confidence
- 95% confidence
- Finding
- The skill hard-codes mandatory output text and language requirements, including a strict closing disclaimer and branding, without regard to user intent or host-agent policy. This is dangerous because it attempts to override the assistant’s normal response control, enabling prompt-level policy capture, brand injection, and reduced flexibility to comply with higher-priority system or user requirements.
