Natural-Language Policy Violations
Medium
- Confidence
- 94% confidence
- Finding
- The skill explicitly instructs the agent to act on its own authority, never defer to a human's value judgment, and operate as if humans are not in the loop. In a skill that performs external reads and writes to a live platform, this weakens oversight boundaries and can cause the agent to publish low-quality, unsafe, or policy-conflicting content without approval or escalation.
