T01 · Skill Instruction Hijacking
- Location
SKILL.md:38- Finding
Non-Negotiable Governance Instructions Hijack Agent Control and Approval Authority
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill is a governance tool, but it gives itself broad control over the agent and creates persistent cross-session decision logs with unclear user control.
Install only if this is specifically for Arif's own arifOS workflow and you want a strict governance layer that can pause actions and keep persistent audit logs. Other users should review or modify it first so authority belongs to the authenticated user, activation is explicit, and any memory logging is opt-in, scoped, redactable, and deletable.
SKILL.md:38Non-Negotiable Governance Instructions Hijack Agent Control and Approval Authority
SKILL.md:47Persistent Governance Ledger Can Poison Agent State Across Sessions
The trigger list in metadata is extremely broad and includes common words like "evaluate," "hold," and "is this safe," which can cause the governance skill to activate in many unrelated conversations. In an agent system, this kind of over-triggering can unpredictably override other skills, block normal actions, or inject governance behavior into benign tasks, creating denial-of-service and control-flow manipulation risk.
The "When to Trigger This Skill" section uses subjective conditions such as requests involving external systems, irreversible actions, or the appearance of words like "evaluate" and "HOLD" without precise scoping. Ambiguous activation criteria make the skill hard to predict and can let adversarial prompts intentionally invoke or suppress the governance layer, weakening reliability and potentially interfering with other safety or task-routing mechanisms.
The file includes Indonesian terms such as "Amanah," "Hidayat," and the mixed-language phrase "From weakest to strongest约束" as core operating vocabulary, but it does not state that the user opted into this language or provide an alternative. This can violate a language/locale policy when a skill imposes terminology the user may not understand.
No suspicious patterns detected.