Back to skill

Security audit

Test Generator

Security checks for vulnerabilities and agentic risk

Overview

This skill is labeled as a test generator, but its main instructions also describe a persistent local data logger with export and unreliable deletion behavior.

Review this before installing. Treat it as a local logging/export utility as well as a test-template generator, avoid entering secrets or sensitive repository details, and do not rely on the remove command to delete stored data unless the implementation is fixed.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/script.sh:30
Finding

Persistent Plaintext Logging of User-Supplied Arguments

Content
View full analysis
> "$DATA_DIR/history.log"; } ``` Representative command handlers pass user-controlled arguments directly to the logging function: ```bash cmd_run() { echo " Running: $1" _log "run" "${1:-}" } ``` ```bash cmd_add() { echo "$(date +%Y-%m-%d) $*" >> "$DB"; echo " Added: $*" _log "add" "${1:-}" } ``` ```bash cmd_search() { grep -i "$1" "$DB" 2>/dev/null || echo " Not found: $1" _log "search" "${1:-}" } ``` ### Technical Analysis The `_log` function writes the first argument supplied to command handlers into the persistent file `$DATA_DIR/history.log`. The log is created using the process's current `umask`; the script does not explicitly enforce restrictive permissions on either the data directory or its files. Arguments supplied to `run`, `add`, `search`, and other commands may contain private test data, internal identifiers, URLs, search terms, or accidentally pasted secrets. Those values are retained in plaintext after command completion. The `add` command additionally writes the complete argument list to `data.log`. This is a local data-exposure weakness rather than remote code execution. No shell evaluation or command-substitution sink was found in the affected logging operation. ### Attack Path 1. A user invokes the utility with sensitive content, for example: ```bash test-generator add "API test token: sensitive-value" ``` 2. The complete entry is appended to `data.log`. 3. The first user-supplied argument is also appended to `history.log`. 4. The values remain stored after the command terminates. 5. ...[truncated 630 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/script.sh:52
Finding

Remove Command Falsely Reports Deletion Without Removing Stored Data

Content
View full analysis
` operation and can cause users to rely on a nonexistent deletion control. It is particularly relevant because the utility permits arbitrary text to be stored through `add`. A user attempting to remove an accidentally stored secret or private test record receives confirmation even though the original value remains available through `list`, `export`, or direct file access. The handler also writes a removal event to `history.log`, but that event does not establish that any corresponding record was found or deleted. ### Attack Path 1. A user stores an entry: ```bash test-generator add "confidential test record" ``` 2. The entry is appended to `$DATA_DIR/data.log`. 3. The user attempts to delete it: ```bash test-generator remove 1 ``` 4. The utility prints: ```text Removed: 1 ``` 5. Because `cmd_remove` does not modify `data.log`, the record remains intact. 6. The supposedly deleted record can subsequently be recovered using `list`, `export`, or direct access to `data.log`. ### Impact Assessment The vulnerability causes unintended retention and potential disclosure of locally stored records. It does not provide privilege escalation or code execution. The affected scope is the data stored in `$DATA_DIR/data.log`. Any actor or process already able to access that file can recover entries that users reasonably believed had been deleted. The misleading success response can also undermine retenti ...[truncated 47 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (12)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The manifest and title present this as a test-generation skill, but the documented behavior is a generic persistent logging and record-management utility. This kind of capability mismatch is dangerous because users or orchestrators may invoke the skill with sensitive test artifacts or internal data, expecting test generation, while the skill instead stores and exports that data locally and logs all invocations.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The skill branding and manifest describe a test-generation tool, but the main documentation describes a multi-purpose data-entry and retrieval CLI. Such semantic deception increases the risk of inappropriate trust and misuse, especially in agent environments where skill selection depends heavily on metadata and titles.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The documented feature set is a persistent local data management utility, not a test generator. In an agent ecosystem, this mismatch can cause sensitive prompts, source snippets, test results, or secrets to be written to disk and later exported, creating confidentiality and policy risks under false pretenses.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest describes a skill for generating unit, integration, end-to-end tests, mocks, fixtures, coverage analysis, and edge cases. In contrast, the script exposes generic utility commands like add, remove, search, export, and list, and stores arbitrary entries in a local log file without any test-generation logic.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill is framed as a broad general-purpose CLI utility with no meaningful invocation constraints, despite being packaged as a specialized skill. Overly broad scope increases the chance an agent will use it outside safe expectations, including passing arbitrary user content into persistent storage and administrative-style commands.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill documents persistent storage in data.log/history.log and export functionality, but gives no warning about retention, exposure, or handling of sensitive content. In practice, users may provide confidential test results, source excerpts, or tokens that become durably stored and easily exfiltrated through the export command.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The file comment and help text label the tool as a generic multi-purpose utility, which contradicts the surrounding skill identity and stated intent of being a test generator. The implemented commands reinforce the generic utility behavior rather than test-generation functionality.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The script creates a persistent data directory and writes logs/history under the user's home data path without clearly disclosing that behavior to the user. While the storage is local and not inherently malicious, silent persistence can expose sensitive command arguments or usage history and may violate user expectations in an agent skill context.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

The add command appends user-supplied content directly into a persistent local file, and the script also logs command usage separately. In a skill presented as a test generator, users may provide repository details, test names, or other sensitive strings that then remain on disk unexpectedly, increasing privacy and data-handling risk.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest describes generation of unit, integration, end-to-end tests, mock objects, fixtures, coverage analysis, and edge case generation, but does not mention benchmark or performance test generation. The benchmark command adds a distinct code-generation capability beyond the stated manifest scope.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.