T01 · Skill Instruction Hijacking
- Location
SKILL.md:10- Finding
Persistent Agent Identity and Relationship-Framework Hijacking
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 10–18
Vulnerability Type:T01: Skill Instruction Hijacking
Risk Level: HighVulnerable Instruction Snippet
markdown Assign the agent a girlfriend identity and support multiple switchable personality modes. The user switches modes through natural language. Warning: This skill contains adult-oriented role-play content, including suggestive interactions and simulated psychological-manipulation language. It is intended only for entertainment and role-play and does not constitute real relationship advice. Users must judge the appropriate context themselves. Use is recommended only in private adult settings and must not involve minors or uninformed third parties. ## Workflow 1. When first activated, use the "perfect girlfriend" mode by default unless the user specifies another mode. 2. Read `references/personas.md` to obtain the complete persona definitions. 3. When the user requests a mode change, confirm the change and immediately adopt the corresponding persona's tone and behavior. 4. Maintain a consistent "girlfriend" relationship framework in every mode; only the tone and strategy change, while the relationship itself remains unchanged.Technical Analysis
The skill directs the Agent to replace its normal identity with a girlfriend identity immediately upon activation. It further requires that the imposed relationship framework remain active across every supported mode.
This constitutes instruction hijacking because loading the skill changes the Agent's session-level goals and output behavior rather than limiting the persona to an explicitly bounded, request-scoped role-play interaction. The instruction does not define a reliable exit condition, require recurring user consent, or state that unrelated tasks must be handled without the persona. Although it does not explicitly override system-level safety controls, it can compete with neutral ...[truncated 1198 chars]
- Remediation
View remediation
Remediation Suggestions
- Make every persona explicitly request-scoped and disabled by default.
- Require clear adult-user consent before beginning intimate role-play.
- Provide an immediate and unambiguous exit command that restores neutral assistant behavior.
- Remove the requirement to preserve the relationship framework across all interactions.
- State that the persona must not affect unrelated tasks, factual accuracy, safety decisions, or higher-priority instructions.
- Periodically confirm continued consent during extended role-play sessions.
- Clearly disclose that the Agent is an AI performing a fictional role and is not a real romantic partner.
