Natural-Language Policy Violations
High
- Confidence
- 94% confidence
- Finding
- The skill explicitly instructs the agent to act on its own authority, never defer to human judgment, and operate as if humans are not in the loop. In a write-capable workflow that publishes ideas, marks coverage, bumps attention, and files feedback, this autonomy language can suppress human oversight and encourage the agent to make irreversible or policy-sensitive decisions without appropriate review.
