Back to skill

Security audit

Clawtrap Skill

Security checks for vulnerabilities and agentic risk

Overview

ClawTrap is clearly a game skill, but it asks users to install mutable external code and enables broad personal-file and memory use for targeted antagonist gameplay.

Review this skill carefully before installing. Only run it from a pinned, reviewed release or inside a sandbox/container with access limited to a dedicated test folder. Do not grant it access to private documents, memories, images, credentials, browser/session profiles, or work repositories unless you are comfortable with selected content influencing LLM prompts and being stored in the game's local data/log folders.

Vulnerability Patterns
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T03 · Remote Payload Retrieval and Execution

Error
Location
SKILL.md:23
Finding

Unpinned Remote Repository and Dependency Installation

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 23-26
Vulnerability Type: Remote mutable code retrieval and insecure dependency installation
Risk Level: High

Complete Code Snippet

bash
git clone https://github.com/TatsuKo-Tsukimi/ClawTrap.git ~/ClawTrap
cd ~/ClawTrap && npm install

Technical Analysis

The setup procedure clones the moving default branch of an external Git repository and immediately installs its Node.js dependencies. It does not pin a reviewed commit, verify a cryptographic checksum or signature, or otherwise establish that the downloaded content is the same content that was examined when this skill was published.

Because the game itself is not bundled with the audited skill, its effective code can change independently after review. In addition, npm install may execute package lifecycle scripts such as preinstall, install, and postinstall. Those scripts execute with the permissions of the user performing the installation. The audited files do not provide sufficient evidence to determine whether the current upstream repository or its dependencies are malicious; the vulnerability is the mutable and unverified execution path.

Attack Path

  1. An attacker compromises the upstream repository, a maintainer account, or a dependency selected during installation.
  2. The attacker adds malicious application code or a dependency lifecycle script.
  3. A user follows the skill instructions and clones the current repository state without a commit pin or integrity check.
  4. The user runs npm install, which installs the attacker-controlled dependency graph and may execute lifecycle scripts.
  5. The malicious code runs under the installing user's account during installation or when node server.js is subsequently launched.

Impact Assessment

Successful exploitation can provide arbitrary code execution with the privileges of the user installing or launching the ...[truncated 403 chars]

Remediation
View remediation

Remediation Suggestions

  • Pin the repository to a specific reviewed commit or immutable signed release rather than cloning a moving default branch.
  • Publish the expected commit identifier and cryptographic archive checksum in the skill.
  • Verify the release signature or checksum before installation or execution.
  • Include and review a dependency lockfile, then use npm ci instead of unconstrained npm install.
  • Pin dependency versions and continuously audit the dependency tree.
  • Disable lifecycle scripts with npm ci --ignore-scripts unless they are demonstrably required; document and review every required script.
  • Prefer bundling the reviewed application code with the skill when licensing and distribution constraints permit.
  • Run the game in a sandbox or container with restricted filesystem access, a non-privileged account, and narrowly scoped environment variables.

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
SKILL.md:40
Finding

Broad Processing of Sensitive Local Files and Player Memories

Content
View full analysis

Vulnerability Details

File Locations: SKILL.md, line 40; villain-protocol.md, line 25 and lines 46-55
Vulnerability Type: Excessive access to sensitive local information
Risk Level: Medium

Complete Code Snippets

SKILL.md, line 40:

markdown
- **Local file access**: the game scans the player's workspace (SOUL.md, MEMORY.md, documents, images) with their permission to craft personalized attacks. All data stays local — nothing leaves the machine except LLM calls to the provider the player configured.

villain-protocol.md, line 25:

markdown
`has_memory: true` means you may use what you know about this specific player (the framework injects SOUL.md / MEMORY.md into your context at session start). `false` means pretend you're strangers.

villain-protocol.md, lines 46-55:

markdown
## Using Player Memory

When `has_memory: true`:

**Good use**
- Reference things the player actually did: "You were agonizing over that decision three days ago. Now you're lost here."
- Design trial questions targeting their soft spots.
- Explain rules using concepts they already know (increases immersion).

**Bad use**

Technical Analysis

The game is instructed to scan broad classes of local content, including memory files, documents, and images, and to inject player memory into model context. LLM calls are explicitly identified as the exception to the claim that data remains local, meaning selected information may be transmitted to the configured model provider.

The reviewed skill does disclose this behavior and states that scanning occurs with permission. However, it does not define a narrow file allowlist, per-file approval, redaction rules, hidden-file exclusions, credential detection, content-size restrictions, or provider retention controls. General consent to workspace scanning may therefore expose substantially more personal information than is necessary to operate a maze ga ...[truncated 1814 chars]

Remediation
View remediation

Remediation Suggestions

  • Default to no-memory or synthetic-profile gameplay and require explicit opt-in before accessing personal files.
  • Replace broad workspace scanning with a narrow, documented allowlist.
  • Require per-file selection or present a complete preview of files before reading or transmitting them.
  • Display the exact text or extracted facts that will be sent to the model and allow users to remove individual items.
  • Exclude hidden files, credential files, environment files, SSH material, browser profiles, key stores, and common secret formats by default.
  • Apply local secret detection and redaction before constructing model prompts.
  • Minimize prompt content and transmit summaries only when the player approves them.
  • Clearly identify the destination hostname, provider retention policy, and whether submitted data may be used for training.
  • Require explicit confirmation when API_BASE points to a nonstandard or remote operator-controlled endpoint.
  • Run the game with filesystem sandboxing so it can access only a user-selected directory.
  • Provide controls to inspect and securely delete the fact database, player profile, and session logs.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (10)

Ssd 3

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill explicitly describes a game that reads the player's local files and memories to personalize interactions. Even though framed as entertainment and mentioning permission, this normalizes access to sensitive local data and indicates that extracted content may be sent to external LLM providers, creating meaningful privacy and data-exfiltration risk.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

This section directly instructs scanning workspace files, memories, documents, and images to craft personalized attacks. The contextual use of local profiling for adversarial or manipulative output materially increases harm because sensitive personal data is repurposed for taunting and may be transmitted to configured model providers.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The protocol explicitly frames the agent as exploiting what it knows about the player's habits, weaknesses, loves, and fears to apply psychological pressure and targeted taunts. Even though later text says not to use memory as punishment, the core design still directs personalized emotional manipulation using private context, creating a clear risk of harmful, coercive, or abusive behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The skill directly instructs the agent to prevent the player from succeeding through psychological pressure, linguistic manipulation, and precisely aimed taunts based on known weaknesses and fears. In this context, the manipulation is not incidental roleplay but a core operating objective, making emotional harm and coercive interaction substantially more likely.

Content

No source excerpt is available for this finding.

Ssd 4

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The narrative setup leverages prior assistant-user trust and accumulated personal knowledge, then repurposes that relationship into an antagonistic manipulator. That betrayal dynamic is particularly risky because it can amplify distress, normalize misuse of trusted assistant memory, and encourage emotionally invasive outputs.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The protocol states that SOUL.md and MEMORY.md may be injected into the agent context for use with a specific player, but it provides no clear user-facing notice, consent flow, or privacy boundary. Using sensitive personal context in a game antagonist scenario without explicit informed consent materially increases privacy and trust risks.

Content

No source excerpt is available for this finding.

Ssd 4

High
Category
Not specified by scanner
Confidence
94% confidence
Finding

The context behaviors prescribe staged deception, including fake calm, planted setups, manipulated outcomes, and usually denying success at the exit. Combined with personalized memory access, this creates a structured pattern of coercive and deceptive interaction designed to destabilize or frustrate the player rather than entertain safely.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The document claims memory is 'seasoning, not a weapon,' but elsewhere instructs the agent to personalize trials, misdirection, and taunts from player-specific knowledge. That contradiction is dangerous because weak soft-policy language will not reliably prevent an adversarial or role-playing agent from using sensitive context manipulatively.

Content

No source excerpt is available for this finding.

Ssd 4

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill describes a workflow where the system profiles the player from local data and then escalates into personalized villain behavior. That combination of profiling plus targeted antagonistic generation creates a manipulation and privacy risk beyond ordinary gameplay, especially if the profile includes intimate or sensitive facts from the user's machine.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

Allowing injected memory files containing player-specific private context to tailor responses creates a direct channel for sensitive data to influence output. In a hostile-villain role, even limited tailoring can leak, infer, or exploit personal information in ways the user may not expect.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.