Back to skill

Security audit

Data Generator

Security checks for vulnerabilities and agentic risk

Overview

This skill is a straightforward training-data generator, but users should know it sends prompts to a configured LLM endpoint and uses an unpinned npx install command.

Install only from a trusted ClawHub/npm environment, consider pinning the installer in your own workflow, and do not feed confidential commands or datasets into this generator unless you trust the API_URL endpoint that will receive the prompt content.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:70
Finding

Unpinned Package Execution Through npx

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 70-75
Vulnerability Type: Unpinned third-party package execution
Risk Level: Medium

Vulnerable Code

markdown
## Installation / 安装

```bash
npx clawhub install data-generator-waai
text

### Technical Analysis

The documented installation procedure invokes `clawhub` through `npx` without specifying a reviewed package version or verifying package integrity. If the package is not already installed locally, `npx` can retrieve the currently published package from its configured registry and immediately execute it.

Consequently, the code executed by this command can change independently of the audited skill. A compromised maintainer account, registry compromise, dependency-confusion condition, or malicious replacement release could cause users following the documentation to execute attacker-controlled code.

This finding concerns the installation instruction itself. The audited Python implementation does not independently retrieve or execute remote source code.

### Attack Path

1. An attacker compromises the package publication account, registry entry, or an upstream component used by the unpinned `clawhub` package.
2. The attacker publishes a malicious release under the package name resolved by `npx`.
3. A user follows the documented installation command.
4. `npx` downloads the currently resolved package release because no trusted version or integrity value is specified.
5. The downloaded CLI or its installation lifecycle code executes with the privileges of the user running the command.
6. The malicious package can access files, environment variables, credentials, and network resources available to that user.

### Impact Assessment

Successful exploitation permits arbitrary code execution under the installing user's account. The affected scope can include the user's project files, home-directory data, accessible environment credentials, an
...[truncated 338 chars]
Remediation
View remediation

Remediation Suggestions

  • Pin clawhub to a specific, reviewed version, for example by using an exact package version rather than resolving the latest release.
  • Record and verify the package integrity hash through an appropriate lockfile or package-manager integrity mechanism.
  • Use an explicitly configured and trusted package registry.
  • Review package lifecycle scripts and the dependency tree before approving version updates.
  • In CI environments, use a pre-reviewed dependency cache or immutable build image rather than downloading executable packages at runtime.
  • Run installation with the least-privileged account possible and avoid exposing unrelated credentials in the installation environment.
  • Establish a controlled update process that reviews new package versions before changing the pinned version.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (10)

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · data_generator.py (reported line 110)May include surrounding context.

python
second = random.randint(0, 59)
        current_time = f"{year}-{month:02d}-{day:02d} {hour:02d}:{minute:02d}:{second:02d}"
        
        return PROMPT_TEMPLATE.format(
            TOOL_NAME=tool_name,
            TOOL_DESCRIPTION=tool_desc,
            USER_COMMAND=user_command,

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill advertises Python-based data generation and the analyzer detected capabilities consistent with environment access, file writing, and network use, but the manifest does not declare any explicit tool scope or permissions boundaries. This creates an avoidable trust gap: consumers cannot tell what resources the skill may access, and an implementation could read sensitive files, exfiltrate data, or write outputs beyond the expected dataset artifacts.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

The installation instruction uses npx clawhub without pinning a specific version, which makes installs dependent on the latest published package state at execution time. If the upstream package is compromised or a breaking/malicious version is released, users may install unreviewed code, leading to supply-chain compromise.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The prompt template is written entirely in Chinese and instructs the model to produce Chinese stylistic filler such as '好的呀~' and '让我来看看~'. This imposes a specific language/locale behavior without any opt-in or alternative language selection, which matches the language policy violation category.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill sends user-provided commands and generated tool context to an HTTP API endpoint, which is an outbound data flow not clearly reflected by the stated purpose of simple local data generation. This creates privacy and trust risk because user inputs may be transmitted to another service, potentially remote if API_URL is changed, without explicit notice or restriction.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The code forwards raw user commands, tool descriptions, and generated context to an LLM API without any explicit warning, consent, or redaction step. If users believe the generator operates locally, this can unintentionally disclose sensitive prompts or embedded data to another service.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
72% confidence
Finding

The file consistently frames the skill in English and Chinese, but it does not explicitly state whether language selection is optional, automatic, or user-controlled. Under the stated policy, locale or language behavior should be offered as a user choice or clearly justified when constrained.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

The README describes generating training data from Excel user instructions and producing JSONL datasets, but it does not warn users about reviewing source data for sensitive, personal, or confidential content before transformation. Because this skill handles user-provided datasets and converts them into reusable training samples, a basic privacy/data-handling warning would help users avoid unintended exposure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

The initializer reads MINIMAX_API_KEY or OPENAI_API_KEY from environment variables, which is access to sensitive credentials under the rule. The docstring notes the source for one key, but there is no clear user warning or disclosure at runtime that the skill accesses credential-bearing environment variables.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

The stated purpose is to generate training data from a tool name and command list, with no mention of credential handling or external model access. Reading MINIMAX_API_KEY and OPENAI_API_KEY introduces access to potentially sensitive environment configuration beyond what the manifest communicates.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.