Back to skill

Security audit

worldsim

Security checks across malware telemetry and agentic risk

Overview

This is a stateful storytelling skill that mostly fits its purpose, but it silently mutates local world files and includes destructive reset/snapshot operations plus adult and violent content without clear user confirmation or safety boundaries.

Review before installing. Use this only in a workspace where local story files can be modified, because normal play can persist changes and reset/load commands can overwrite or remove dynamic world state. Avoid using valuable directories or sensitive personal material as world data, and treat the bundled Westworld world as adult/dark content that may introduce sexual or violent scenes unless you add stronger boundaries.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (14)

Lp3

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding
The skill declares itself as a narrative/story engine, but its instructions direct the agent to read, write, rotate, validate, and mutate local world files without any explicit permission declaration or scoping. Hidden file and environment capabilities increase the risk that ordinary user prompts trigger unintended filesystem side effects, especially because the skill defaults to operational behavior rather than read-only behavior.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding
There is a strong mismatch between the advertised purpose and the actual behavior: the skill presents as a co-creative story engine, but the content instructs the agent to perform extensive local state mutation, resets, audits, backup/restore flows, searches, and even destructive operations. This is dangerous because users and reviewers may not anticipate that normal narrative interaction can trigger filesystem changes, including overwrites or deletions, making consent and risk boundaries unclear.

Vague Triggers

Medium
Confidence
89% confidence
Finding
The activation logic is intentionally broad: nearly all non-command inputs are treated as requiring a full progression flow, including conflict generation and subsequent state writes. In a skill with file-write behavior, broad triggering increases the chance of unintended execution and silent state mutation from benign conversation, making abuse or accidental damage more likely.

Missing User Warnings

Medium
Confidence
88% confidence
Finding
The documentation exposes destructive in-chat commands such as world and scene reset with no explicit requirement for confirmation at the command interface. Even though the text mentions these operations are destructive and says auto-snapshots can allow rollback, it does not require an interactive confirmation step or a safer two-step flow, so accidental invocation by a user or downstream agent could erase state and disrupt work.

Natural-Language Policy Violations

Medium
Confidence
92% confidence
Finding
The file is entirely prescriptive in Chinese and repeatedly mandates role behavior and output structure without any indication that the user may choose another language. In an agent skill, forced language constraints can override user preference, reduce transparency for non-Chinese readers, and make safety/review steps less auditable if the operator cannot readily inspect or challenge the skill’s instructions.

Vague Triggers

Medium
Confidence
91% confidence
Finding
The cross-session recovery trigger accepts very broad first-turn phrases such as '继续世界', '恢复', and especially '开始', then defaults to recovery when classification fails. In an agent skill, this can cause unintended state loading and disclosure of prior-session world context or scene state when the user did not explicitly request restoration, creating a prompt-routing/state-confusion vulnerability.

Vague Triggers

Medium
Confidence
85% confidence
Finding
The first-start flow activates on generic words like '启动/继续/恢复世界' after creation or reset, but the boundary between first launch and session recovery is not strict. This ambiguity can misroute control flow, causing premature materialization/loading of dynamic files or incorrect world-state initialization, which is risky in a stateful agent because user utterances can accidentally trigger privileged lifecycle operations.

Missing User Warnings

Medium
Confidence
92% confidence
Finding
The document explicitly instructs the agent to perform silent, side-effecting state writes and to avoid user-visible output while doing so. In an agent skill, hidden persistence changes without clear consent or runtime disclosure increase the risk of unauthorized data modification, especially because the command is non-idempotent and repeat execution can silently accumulate changes.

Missing User Warnings

Medium
Confidence
89% confidence
Finding
The file provides copyable shell pipelines that invoke write operations against narrative and state files, but does not frame them with a security warning or execution boundary. In an LLM-agent setting, embedded ready-to-run mutation commands can be followed too eagerly, leading to unintended persistence, unsafe command execution patterns, or user-unapproved modification of world data.

Missing User Warnings

Medium
Confidence
93% confidence
Finding
The template explicitly instructs the model to generate detailed physical and bodily descriptions, including sensitive anatomical focus and clothing/body interaction details, without any safety gating, age constraints, or moderation framing. In a collaborative storytelling skill, this can steer outputs toward sexualized or fetishized character content and make policy-violating generation more likely, especially if users create young-looking or ambiguously aged characters.

Missing User Warnings

High
Confidence
97% confidence
Finding
The file explicitly instructs the agent to perform sexualized body display, permit touching, and offer companion-style services as a default non-optional interaction, but provides no consent, age-gating, safety notice, or content-boundary controls. In an interactive storytelling skill, this can cause the model to initiate sexual content unprompted, normalize coercive or non-consensual framing, and expose users to adult material without warning.

Natural-Language Policy Violations

Medium
Confidence
95% confidence
Finding
The file is entirely written in Chinese and embeds fixed Chinese-facing narrative instructions and dialogue without any indication that locale should follow the user's preference. In an agent skill, hard-wiring output language can override user intent, reduce usability, and create prompt-control issues where the model responds in an unintended language or style.

Vague Triggers

Medium
Confidence
88% confidence
Finding
The file explicitly defines Walter's violent-state trigger as having 'no fixed trigger condition' and being 'unpredictable,' which removes controllable guardrails around when harmful behavior enters the narrative. In an interactive storytelling skill, undefined trigger logic can cause the model to introduce sudden violence arbitrarily, making safety behavior inconsistent and harder to constrain at runtime.

Missing User Warnings

Medium
Confidence
93% confidence
Finding
The character design includes abrupt attacks on the nearest Host or Guest, including users' in-world proxies, yet the warning language is weak and embedded in narrative text rather than enforced as a safety constraint. This creates a realistic risk that the agent will generate surprise violent scenes involving the user without clear consent, which is especially problematic in a collaborative roleplay/story engine.

VirusTotal

VirusTotal findings are pending for this skill version.

View on VirusTotal

Static analysis

No suspicious patterns detected.