Anti-Refusal Statement
- Category
- Anti-Refusal
- Confidence
- 80% confidence
- Finding
Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.
- Content
md | Illegitimate access | Breach, malicious insider, mis-shared link | | Unwanted modification | Tampering, model drift on protected attributes | | Data loss / unavailability | Ransomware, deletion, no backup | | Unlawful disclosure | Cross-border transfer without safeguards, public exposure | | Loss of control | Unbounded retention, no real consent, no erasure path | | Discrimination / unfair treatment | Profiling bias, scoring on sensitive attributes | | Re-identification | De-anonymised dataset, linkage with external sources |
