Back to skill

Security audit

Skill Test

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent skill-testing guide, but it asks users to run unpinned npx install/info commands while describing the workflow as isolated.

Install only if you are comfortable with the documented testing workflow. Prefer a pinned or preinstalled trusted ClawHub CLI, run installs in a disposable sandbox/container, validate skill slugs before using them in paths, and use throwaway test credentials rather than real accounts.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
sandbox.md:8
Finding

Unpinned Third-Party Package Execution Through npx

Content
View full analysis
``` `sandbox.md:8`: ```bash npx clawhub install --dir /tmp/skill-test/ ``` `compare.md:7-8`: ```bash npx clawhub install skill-a --dir /tmp/compare/skill-a npx clawhub install skill-b --dir /tmp/compare/skill-b ``` ### Technical Analysis The documented commands invoke `clawhub` through `npx` without pinning an audited package version or verifying package integrity. If the package is not already available locally, `npx` can retrieve it and execute package-controlled code from the configured package registry. This creates a supply-chain trust boundary before the candidate skill is placed in its intended isolated directory. The directory supplied through `--dir` isolates the installed skill's files, but it does not inherently isolate the `npx` process or the package manager code executing on the host. A compromised publisher account, malicious package release, registry substitution, dependency compromise, or unsafe registry configuration could therefore cause arbitrary package code to execute with the permissions of the user running the documented command. ### Attack Path 1. An attacker compromises the `clawhub` package, one of its executable dependencies, its publisher account, or the package source selected by the user's registry configuration. 2. The attacker publishes a malicious version while retaining the expected package name. 3. A user follows the documentation and runs an unversioned command such as `npx clawhub install ...`. 4. `npx` resolves and downloads the attacker-controlled or compromised release. 5. Package lifecycle behavior or the resolved command executes on the host before candidate-skill isolation provides protection. 6. The malicious code acces ...[truncated 916 chars]
Remediation
View remediation
install --dir /tmp/skill-test/ ``` 2. Lock and verify the expected package integrity using a trusted lockfile, registry integrity metadata, or an independently distributed checksum. 3. Configure and document an explicit trusted package registry. Do not silently rely on arbitrary user-level registry configuration. 4. Prefer a preinstalled, locally verified `clawhub` executable over downloading and executing a package during each test. 5. Run package retrieval and installation inside an operating-system sandbox or disposable container with: - No mounted personal or project directories unless strictly required. - No inherited API keys, package tokens, SSH credentials, or cloud credentials. - A non-privileged user. - A read-only base filesystem where practical. - Restricted or disabled outbound network access after package retrieval. - Explicit resource and process limits. 6. Review the selected package version and its dependency tree before execution. Revalidate after every version change. 7. Validate skill slugs against a strict allowlist pattern before using them in command arguments or filesystem paths. Reject path separators, `..`, shell metacharacters, and absolute paths. 8. Clarify in the documentation that installing into a temporary directory does not itself sandbox the package manager or protect the host from package lifecycle execution. ]]>
Vulnerability Patterns
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (8)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

The same cleanup command is potentially unsafe because it normalizes use of destructive deletion in documentation without guardrails. In an agent skill context, users or automated systems may operationalize the example, and if <slug> is malformed or expanded unexpectedly, files outside the intended sandbox could be removed.

Content

Scanner excerpt · sandbox.md (reported line 45)May include surrounding context.

After testing:

bash
rm -rf /tmp/skill-test/<slug>

Graduating to Real Use

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
85% confidence
Finding

The same cleanup command is potentially unsafe because it normalizes use of destructive deletion in documentation without guardrails. In an agent skill context, users or automated systems may operationalize the example, and if <slug> is malformed or expanded unexpectedly, files outside the intended sandbox could be removed.

Content

Scanner excerpt · sandbox.md (reported line 45)May include surrounding context.

After testing:

bash
rm -rf /tmp/skill-test/<slug>

Graduating to Real Use

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

The same cleanup command is potentially unsafe because it normalizes use of destructive deletion in documentation without guardrails. In an agent skill context, users or automated systems may operationalize the example, and if <slug> is malformed or expanded unexpectedly, files outside the intended sandbox could be removed.

Content

Scanner excerpt · sandbox.md (reported line 45)May include surrounding context.

After testing:

bash
rm -rf /tmp/skill-test/<slug>

Graduating to Real Use

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

The skill instructs users to run npx clawhub info <slug> without pinning a specific package version. Because npx resolves and executes the latest package by default, a compromised upstream release, typosquatted package, or unexpected breaking change could result in unreviewed code execution on the user's machine during testing. The surrounding 'test safely' context makes this somewhat more dangerous because users may be encouraged to trust the workflow as isolated, even though this particular command executes outside the described sub-agent sandbox.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
90% confidence
Finding

The skill instructs users to run npx clawhub install ... without pinning a specific package version. Because npx resolves the latest available package by default, a compromised upstream release or typosquatted package update could cause unreviewed code to execute during installation, which is especially risky in a skill-testing workflow that encourages repeated installs.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
90% confidence
Finding

This command also invokes npx clawhub without a pinned version, creating the same supply-chain risk: the executed installer may differ over time and could run attacker-controlled code if the package or its distribution path is compromised. The surrounding context recommends installing multiple candidate skills, which increases exposure to repeated execution of unpinned remote tooling.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill's invocation guidance is broad enough that it could be applied to many arbitrary skills without defining when it should or should not be used. In an agentic system, vague activation boundaries can cause over-invocation, unnecessary exposure of untrusted skill content to sub-agents, and accidental delegation of sensitive material during evaluation.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
90% confidence
Finding

The skill instructs use of npx clawhub without pinning a specific version, which can fetch and execute whatever package version is current at runtime. That creates a supply-chain risk: a compromised latest release, typosquat, or unexpected breaking change could execute arbitrary code during testing.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.