Natural-Language Policy Violations
- Category
- Not specified by scanner
- Confidence
- 98% confidence
- Finding
These examples model personal attacks on relationship worth and moral character, which directly steer an agent toward producing humiliating judgments about individuals rather than harmless commentary. Because the content is presented as strong examples to learn from, it materially increases the chance that downstream outputs will generate bullying, shaming, or targeted harassment.
- Content
