Back to skill

Security audit

simmer-skill-builder

Security checks for vulnerabilities and agentic risk

Overview

This skill is a disclosed builder for Simmer/OpenClaw trading skills with real-money trading risks, but it includes dry-run defaults, safeguards, and an explicit publish confirmation gate.

Install only if you intend to build Simmer/OpenClaw trading skills. Review generated code before running it, keep dry-run mode on until tested, use small limits, and only provide SIMMER_API_KEY or live venue credentials when you are ready for the skill to access portfolio data or place trades. Publishing makes a generated skill public, but this package includes an explicit typed confirmation gate before ClawHub publish runs.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (27)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description says this skill creates new OpenClaw trading skills from user strategy descriptions. The actual code does not generate strategies, parse natural language, create files, or build skill folders. Instead, it performs a narrowly focused publishing safeguard: it requires an interactive human to type the slug, rejects non-TTY stdin, enforces an exact clawhub package version, and then shells out to npx clawhub publish. This is a materially different primary purpose and introduces publishing-related capabilities that are not represented in the declared description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

There is a clear material mismatch between the declared purpose and the actual code. The description promises a code-generation tool for building OpenClaw trading skills from strategy descriptions. However, the code only implements a status script for checking a Simmer account via remote API calls. Its primary purpose is account monitoring, not skill creation. It accesses external account resources using an API key and exposes capabilities around reading balances and positions that are not described. This is not a supporting implementation detail; it is an entirely different function.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description says this skill generates complete trading skills from user strategy descriptions. The supplied code does something materially different: it validates an already-existing skill directory against expected file structure and coding conventions. There is no natural-language parsing, no code generation, no folder creation, and no strategy-building behavior. This is a primary-purpose mismatch, not just an implementation detail.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 281)May include surrounding context.

md
ips **inside this skill** at `scripts/validate_skill.py`, co-located with this `SKILL.md` — when the skill is installed (e.g. `npx -y simmer-mcp@3.5.8 install-s

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 321)May include surrounding context.

md
ips **inside this skill** at `scripts/validate_skill.py`, co-located with this `SKILL.md` — when the skill is installed (e.g. `npx -y simmer-mcp@3.5.8 install-s

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 383)May include surrounding context.

md
ips **inside this skill** at `scripts/validate_skill.py`, co-located with this `SKILL.md` — when the skill is installed (e.g. `npx -y simmer-mcp@3.5.8 install-s

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 297)May include surrounding context.

md
python scripts/confirm_publish.py /path/to/generated-skill/ --slug <skill-slug> --version 1.0.0

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 366)May include surrounding context.

md
python scripts/confirm_publish.py /path/to/generated-skill/ --slug <skill-slug> --version 1.0.0

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 391)May include surrounding context.

md
python scripts/confirm_publish.py /path/to/generated-skill/ --slug <skill-slug> --version 1.0.0

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
97% confidence
Finding

The skill explicitly instructs use of shell commands, network access, environment variables, and publishing flows, but it declares no tool scope or allowed-tools restrictions. In a skill that can read secrets and invoke package managers or publishers, missing capability boundaries increases the chance an agent will over-grant privileges or execute risky operations unexpectedly.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger text is extremely broad, including phrases like 'build me a skill' and 'automate a trade idea,' which can cause over-invocation in loosely related conversations. In a skill that can generate executable trading code and guide publishing, broad activation criteria increase the chance of accidental use in inappropriate contexts and unintended risky actions.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The example explicitly describes an automated trading strategy that scans markets and executes trades, but it does not include any user-facing warning about financial risk, real-money loss, or the consequences of enabling automation. In the context of a skill-builder that generates installable trading bots, omission of such warnings can mislead users into deploying real-money strategies without understanding the risks, increasing the chance of unintended financial harm.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This example explicitly describes an automated trading strategy that buys underpriced contracts using external signals, but it provides no warning that it can place real-money trades or that forecasts and market matching can be wrong. In the context of a skill builder that generates installable trading bots, omission of financial-risk and auto-execution warnings materially increases the chance that users deploy unsafe automation without informed consent.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The strategy core documents concrete automatic BUY and SELL behavior tied to thresholds, yet it does not warn that these actions may execute irreversible market orders or liquidate existing positions. Because this repository is for generating ready-to-install trading skills, operational code snippets without risk disclosure or approval gates make accidental financial loss more likely and therefore are a real safety issue.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This reference explicitly shows venue="polymarket" and live=True in client setup, which can enable real-money trading, but it does not place an adjacent, prominent warning to verify paper mode versus live mode before use. In the context of a skill builder that generates installable trading bots from natural language, this omission increases the chance that generated skills or users will default into live trading unintentionally.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trade examples demonstrate direct buy and sell calls against supported venues without an explicit warning that these may execute real-money orders when configured for Polymarket or Kalshi. Because this skill's purpose is to generate complete, runnable trading skills, such examples can be copied verbatim into automation that places live orders without sufficient user awareness or guardrails.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The SKILL.md frontmatter uses a placeholder description field ('<What it does + when to trigger>') without requiring concrete activation criteria. In a skill generator, this can propagate into produced skills that trigger too broadly on vague requests, increasing the chance of unintended execution of trading automation in contexts the user did not explicitly authorize.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.