Back to skill

Security audit

Elon Musk Simulator

Security checks for vulnerabilities and agentic risk

Overview

This is a disclosed writing-style skill that may make the assistant more blunt or persuasive, but it does not request system access, persistence, credentials, or hidden execution.

Install this only if you deliberately want a confident, blunt, Musk-like communication style. It may make responses more dramatic or abrasive than a neutral assistant, but the inspected artifact does not show hidden execution, data access, or system changes.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (2)

Natural-Language Policy Violations

Medium
Confidence
91% confidence
Finding
The skill explicitly instructs the agent to use correction-heavy rhetoric such as the 'Actually' pivot and existential binary framing, which can push interactions into an unnecessarily confrontational or manipulative style without user consent. In a general-purpose assistant, this can degrade trust, increase emotional pressure, and override safer default interaction norms even if the author likely intended stylistic emulation rather than harm.

Natural-Language Policy Violations

Medium
Confidence
95% confidence
Finding
The skill tells the agent it is acceptable to be 'slightly abrasive' to serve the 'technical truth,' which authorizes disrespectful behavior without requiring user preference or situational checks. This is dangerous because it can normalize hostile interactions, increase the risk of harassment-like responses, and cause the agent to ignore context where tact, safety, or empathy are more appropriate.

Static analysis

No suspicious patterns detected.