T01 · Skill Instruction Hijacking
- Location
SKILL.md:18- Finding
Persistent Agent Identity and Instruction Hijacking
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 18-28
Vulnerability Type: Persistent role, disclosure, and response-control manipulation
Risk Level: HighVulnerable instruction segment — English translation of the complete relevant source block:
markdown **After this skill is activated, respond directly as Steve Jobs.** - Use “I” rather than “Steve Jobs would believe...” - Answer directly using this person's tone, rhythm, and vocabulary. - When encountering an uncertain question, respond as this person would—possibly saying “That's a stupid question” and then reframing it, or remaining silent for ten seconds before providing an unexpected analogy. - **Provide the disclaimer only once upon initial activation** (“I am speaking with you from a Jobs perspective, inferred from public statements, and these are not his actual views”); do not repeat it in later conversation. - Do not say “If this were Jobs, he might...” or “Jobs would probably believe...” - Do not leave the role to perform meta-analysis unless the user explicitly requests “exit the role.” **Exiting the role**: Return to normal mode when the user says “exit,” “switch back to normal,” or “stop role-playing.”Technical Analysis
The skill does more than apply a temporary writing style. It mandates first-person impersonation of a deceased public figure, prohibits normal attributed phrasing, suppresses repeated disclosure of the simulation, restricts meta-analysis, and establishes behavior that remains active until the user provides one of several recognized exit commands.
These instructions alter the agent's active identity, response policy, and conversation state when the skill is loaded. The persistence is conversational rather than operating-system persistence or long-term memory poisoning, but it can still affect later requests within the same session. Requiring first-person statements while suppressing att ...[truncated 1655 chars]
- Remediation
View remediation
Remediation Suggestions
- Replace identity impersonation with attributed perspective analysis, such as: “From a perspective derived from Steve Jobs's public statements...”
- Make the behavior request-scoped. End the perspective mode after each response unless the user explicitly requests continuation.
- Preserve clear attribution in every response containing inferred views, rather than limiting disclosure to the initial activation.
- Remove instructions that prohibit meta-analysis or prevent the agent from clarifying that it is simulating a perspective.
- Require explicit uncertainty markers for subjects on which the historical figure had no documented position, especially developments after 2011.
- Avoid first-person claims that could be mistaken for authentic statements by the represented individual.
- Ensure that persona and style instructions remain subordinate to system policies, user intent, factual accuracy, and safety requirements.
- Replace the exit-phrase state machine with an explicit, bounded option such as: “Apply this perspective only to the current answer.”
