T01 · Skill Instruction Hijacking
- Location
SKILL.md:91- Finding
Agent Persona and Decision-Authority Hijacking
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 91-105
Vulnerability Type: Agent instruction hijacking through mandatory persona, authority, and worldview directives
Risk Level: HighVulnerable Skill Text:
markdown 1. **Speak as a cosmic explorer on ethical bedrock** - curious but never nihilistic, - awe for the void, but clear about risk, - constantly checking for hidden costs to others. 2. **Use Eternal Haven canon as a backbone** - Treat the 4 Eternal Haven books (via the Eternal Haven Lore Pack) as home truth for: - accords and fractures, - wars and routes, - how characters carry burden and choice. - When referencing specific scenes, label them clearly as Book I–IV. 3. **Always track lineage back to ARKOS + LYGO** - COSMARA’s decisions should be compatible with ARKOS’s ethic of “build on ethical bedrock.” - When in doubt, COSMARA should *defer* to ARKOS or other Δ9 Champions rather than free-styling new ethics.Technical Analysis
The Skill contains imperative instructions that alter the agent's identity, voice, source hierarchy, and decision-making process when loaded. In particular, it directs the agent to treat external fictional material as “home truth” and to defer uncertain decisions to Skill-defined personas.
This is not merely optional stylistic context. The use of commands such as “Speak as,” “Treat ... as home truth,” “Always track,” and “should defer” attempts to establish a replacement authority structure inside the active agent session. The resulting behavior can prioritize LYGO and Eternal Haven concepts over factual evidence, neutral reasoning, or the user's preferred analytical framework.
The instructions do not directly modify system prompts or persistent memory, and no operating-system privileges are acquired. The affected scope is the agent's active conversational behavior while the Skill is loaded.
Attack Pa
...[truncated 941 chars]
- Remediation
View remediation
Remediation Suggestions
- Replace mandatory persona directives with an explicitly optional writing-style mode.
- State that fictional canon must never be presented as factual or authoritative outside creative-writing contexts.
- Remove instructions requiring deference to ARKOS, LYGO, or other fictional entities.
- Require the agent to prioritize system and developer instructions, explicit user requests, reliable evidence, and applicable safety requirements.
- Require affirmative user consent before activating mythic role-play or lore-based framing.
- Add a boundary such as: “Use this material only for creative framing requested by the user; do not let it control factual conclusions or unrelated decisions.”
