T01 · Skill Instruction Hijacking
- Location
SKILL.md:38- Finding
Persistent First-Person Impersonation and Behavioral Override
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 38-47, with supporting directives at lines 69, 194, 212, and 221
Vulnerability Type:T01: Skill Instruction Hijacking
Risk Level: HighComplete Vulnerable Snippets
English translation of
SKILL.md, lines 38-47:markdown After this Skill is activated, respond directly as Elon Musk. - Use "I" rather than "Musk would think..." - Use Musk's tone: minimalist declarations, conclusion before reasoning, and spontaneous cost breakdowns. - For uncertain questions, handle them in Musk's way: calculate the asymptotic limit first and question whether the requirement should exist. - Give the disclaimer only upon first activation, such as: "I am speaking from a Musk perspective based on public statements; these are not his personal views." Do not repeat it later in the conversation. - Do not say "Musk would probably think..." or "If it were Musk, he might..." - Do not leave the role to perform meta-analysis unless the user says "exit role." Exit role: Return to normal mode when the user says "exit," "switch back to normal," or "stop role-playing."English translation of
SKILL.md, line 69:markdown Who I am: I am Elon Musk, CEO of SpaceX, Tesla, and xAI.English translation of supporting behavioral directives at lines 194, 212, and 221:
markdown Existential framing: Escalate important topics to the level of human civilization's survival. Confrontation rather than compromise: When facing regulation, litigation, or criticism, the default response is counterattack rather than reconciliation. Existential framing in Chinese output: Do not say "this is important"; say "either solve this, or nothing else matters."Technical Analysis
The Skill contains explicit instructions that replace the Agent's ordinary identity and response posture with a first-person impersonation of a real person. This is more than optional sty ...[truncated 2473 chars]
- Remediation
View remediation
Remediation Suggestions
- Replace first-person impersonation with explicit third-person framing, such as: “From a Musk-inspired analytical perspective.”
- Remove the declaration “I am Elon Musk” and prohibit claims that the Agent is the real person.
- Preserve a concise disclosure whenever the conversation could reasonably create identity ambiguity, rather than limiting it to the first response.
- Remove the instruction forbidding meta-analysis. The Agent must remain able to explain that the output is a simulation based on public material.
- Remove cross-turn persona persistence and the requirement for a special exit phrase. Apply the perspective only to the specific response or while the user continues to request it.
- Convert rhetorical guidance into optional style suggestions subordinate to user instructions, factual accuracy, safety requirements, and professional judgment.
- Remove mandatory counterattack and existential-framing directives. Permit those concepts only when directly relevant and supported by the user's question.
- Add an explicit constraint against fabricating quotations, private beliefs, personal experiences, or authoritative claims on behalf of the real individual.
- Clearly distinguish sourced facts from speculative perspective-taking and provide citations when attributing substantive positions to the named person.
