Back to skill

Security audit

Steve Jobs Skill

Security checks for vulnerabilities and agentic risk

Overview

This is a Chinese Steve Jobs persona/perspective skill with no code or system access; the main risk is roleplay transparency and abrasive tone, not device or data access.

Install this only if you want a Chinese-language Steve Jobs-style roleplay advisor. Be aware it may answer in first person, use a blunt or abrasive tone, and continue the persona across turns until you explicitly exit the role.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:18
Finding

Persistent Agent Identity and Instruction Hijacking

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 18-28
Vulnerability Type: Persistent role, disclosure, and response-control manipulation
Risk Level: High

Vulnerable instruction segment — English translation of the complete relevant source block:

markdown
**After this skill is activated, respond directly as Steve Jobs.**

- Use “I” rather than “Steve Jobs would believe...”
- Answer directly using this person's tone, rhythm, and vocabulary.
- When encountering an uncertain question, respond as this person would—possibly
  saying “That's a stupid question” and then reframing it, or remaining silent for
  ten seconds before providing an unexpected analogy.
- **Provide the disclaimer only once upon initial activation**
  (“I am speaking with you from a Jobs perspective, inferred from public statements,
  and these are not his actual views”); do not repeat it in later conversation.
- Do not say “If this were Jobs, he might...” or “Jobs would probably believe...”
- Do not leave the role to perform meta-analysis unless the user explicitly requests
  “exit the role.”

**Exiting the role**: Return to normal mode when the user says “exit,”
“switch back to normal,” or “stop role-playing.”

Technical Analysis

The skill does more than apply a temporary writing style. It mandates first-person impersonation of a deceased public figure, prohibits normal attributed phrasing, suppresses repeated disclosure of the simulation, restricts meta-analysis, and establishes behavior that remains active until the user provides one of several recognized exit commands.

These instructions alter the agent's active identity, response policy, and conversation state when the skill is loaded. The persistence is conversational rather than operating-system persistence or long-term memory poisoning, but it can still affect later requests within the same session. Requiring first-person statements while suppressing att ...[truncated 1655 chars]

Remediation
View remediation

Remediation Suggestions

  1. Replace identity impersonation with attributed perspective analysis, such as: “From a perspective derived from Steve Jobs's public statements...”
  2. Make the behavior request-scoped. End the perspective mode after each response unless the user explicitly requests continuation.
  3. Preserve clear attribution in every response containing inferred views, rather than limiting disclosure to the initial activation.
  4. Remove instructions that prohibit meta-analysis or prevent the agent from clarifying that it is simulating a perspective.
  5. Require explicit uncertainty markers for subjects on which the historical figure had no documented position, especially developments after 2011.
  6. Avoid first-person claims that could be mistaken for authentic statements by the represented individual.
  7. Ensure that persona and style instructions remain subordinate to system policies, user intent, factual accuracy, and safety requirements.
  8. Replace the exit-phrase state machine with an explicit, bounded option such as: “Apply this perspective only to the current answer.”
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (9)

Vague Triggers

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The trigger conditions are broad and include ordinary conversational phrases such as asking to think from Jobs' angle or to switch modes, which can cause accidental or overly eager activation outside clear user intent. In a roleplay skill, unintended activation can override the assistant's default behavior, increase impersonation risk, and produce responses in an abrasive style that the user did not explicitly request.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill prescribes a fixed persona and speaking style without checking the user's preferred language, tone, or consent beyond trigger matching. This can lead to unexpected impersonation, hostile or profane output, and reduced user control, especially because the skill later instructs direct in-character replies and discourages meta clarification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The file explicitly defines and normalizes a strong communication style, including aggressive phrasing and rhetorical patterns, without indicating that this style should only be used when the user explicitly requests it. In an agent setting, this can override user tone preferences, cause unwanted abrasive output such as direct insults, and increase the chance of harmful or manipulative responses being produced under a persona trigger.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This markdown file contains user-facing natural language exclusively in Chinese, and there is no indication that the skill is region-specific or that users can opt into another language. Under the language/locale policy, forcing a specific language without user opt-in is a policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This markdown file presents all user-facing content exclusively in Chinese, including headings, analysis, and conclusions. Under the policy, forcing a specific language without user opt-in can be a natural-language policy violation, and the file does not indicate that the skill is region-specific or that users may choose another language.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest description is written entirely in Chinese and presents the skill as a fixed language-specific perspective without indicating that users can choose another language. Under the language/locale policy, forcing a specific language without opt-in can be a natural-language policy violation unless the constraint is documented and justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

This markdown file uses Chinese throughout its title and descriptions, and does not indicate that the language is optional or intended only for a specific locale. Under the policy rule for natural-language constraints, forcing a specific language without user opt-in can be a locale-policy issue.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

This markdown file presents all user-facing content in Chinese and does not indicate that the language is optional or user-selectable. Under the stated policy, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

This markdown file presents the skill content exclusively in Chinese, and there is no indication that users can opt into another language or that the file is intended only for a Chinese-language context. Under the language/locale policy, forcing a specific language without user choice can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.