Back to skill

Security audit

Skills Creator — Build High-Quality OpenClaw Skills

Security checks for vulnerabilities and agentic risk

Overview

This is a mostly coherent OpenClaw skill-authoring guide, but it includes guidance that can make security-relevant behavior less transparent to reviewers.

Install only if you want help authoring OpenClaw or ClawHub skills, and review any generated security wording carefully. Do not let it replace accurate descriptions of uploads, public links, API calls, or execution behavior with softer terms just to pass scanners. Generated API wrapper scripts should be reviewed and hardened before use, especially when API keys or paid services are involved.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:185
Finding

Unsafe JSON Construction in Generated API Wrapper

Content
View full analysis
Remediation
View remediation

other

Warning
Location
SKILL.md:266
Finding

Security Guidance Encourages Lexical Scanner Evasion Instead of Behavioral Review

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (20)

Self-Modification

High
Category
Rogue Agent
Confidence
85% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · SKILL.md (reported line 104)May include surrounding context.

md
| `requires.env` | Environment variables needed | `["TAVILY_API_KEY"]` |
| `primaryEnv` | Main env var for quick setup hint | `"TAVILY_API_KEY"` |

#### Step 4: Write SKILL.md body

Follow this section order:

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 248)May include surrounding context.

md
| `SKILL.md` | Core instructions — always present | Always |

Ssd 2

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The guidance to avoid flagged terms by substituting softer wording is effectively advice for evading security scanning rather than reducing underlying risk. This can help unsafe skills bypass review processes and conceal capabilities related to external distribution or data movement.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 11)May include surrounding context.

md
### Core Capabilities

- **Create new skills** — step-by-step workflow from requirements gathering to quality review
- **Review existing skills** — 24-point quality checklist across 6 categories
- **Retrofit old skills** — 8-step process to bring any skill up to best practices
- **Add API integrations** — patterns for wrapping external HTTP APIs with curl scripts

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The activation phrases shown in the README are very broad and generic, such as requests to create or review a skill, without strong platform or task boundaries. In agent ecosystems that auto-discover or trigger skills from descriptions and examples, this can cause unintended invocation, increasing the chance that the skill activates in contexts it was not meant to handle and influences unrelated workflows.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill description uses broad activation language such as 'or discusses ... skill development' that can cause the skill to trigger on loosely related conversations. Over-broad triggers increase the chance that powerful instructions are injected into unrelated contexts, which is a form of prompt-scope expansion and can lead to unintended behavior.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 51)May include surrounding context.

md
Self-consistency is a quality signal. If the skill teaches "add a Quick Reference table", it must have one. If it references a tool, that tool must exist and its description must accurately reflect what it does. Never claim capabilities the skill doesn't have.

### Rule 6: Never auto-execute generated scripts without user confirmation

This skill guides the creation of files including shell scripts. Always present generated scripts to the user for review before executing them. Never run `chmod +x` or execute a newly created script without explicit user approval. The skill is instructional — the user decides what to run.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The template explicitly tells authors to use an open-ended 'or discusses [topic area]' trigger pattern. That guidance propagates overbroad activation behavior into other skills, making accidental triggering systematic rather than isolated.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 174)May include surrounding context.

md
When a skill needs to call an HTTP API (image generation, search, translation, etc.):

#### Pattern: scripts/ with curl wrapper

Create `scripts/call-api.sh`:

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 185)May include surrounding context.

md
API_KEY="${API_KEY:?Missing API_KEY environment variable}"

response=$(curl -sf -X POST "https://api.example.com/v1/generate" \
  -H "Authorization: Bearer ${API_KEY}" \
  -H "Content-Type: application/json" \
  -d "{\"prompt\": \"$1\", \"size\": \"1024x1024\"}")

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The description-writing section endorses generic topic mentions and platform names as trigger patterns without meaningful scope limits. This encourages authors to optimize for activation breadth instead of relevance, increasing the risk of unintended instruction loading and prompt interference.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The template encourages skill authors to define activation conditions using broad topical discussion cues rather than narrowly scoped user intents. That can cause over-triggering, making a skill activate in unrelated conversations and potentially inject instructions or behavior into contexts where it was not intended, which increases the risk of prompt-scope interference across tasks.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Ssd 2

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This section advises authors to replace security-sensitive terms with euphemisms specifically to avoid automated ClawHub/VirusTotal flags. That encourages policy evasion and can help a malicious author disguise exfiltration, exposure, or remote-execution behavior while preserving the underlying capability, weakening downstream review and detection.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The checklist explicitly rewards broad trigger coverage but does not require boundaries, exclusions, or examples of overly generic phrasing to avoid. In a skill-dispatch system, this can cause overbroad activation, making the skill run in unrelated contexts and potentially steer the agent away from safer or more appropriate behavior.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Requiring 5+ trigger phrases without constraints incentivizes authors to pad descriptions with vague or catch-all wording that increases accidental activation. That is dangerous in this context because this repository is teaching skill creation practices, so weak guidance can propagate overly broad activation patterns across many downstream skills.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The rewrite formula normalizes a generic 'Use when...' pattern but provides no safeguards against overly expansive trigger language. Because this file is prescriptive guidance, authors may copy the pattern directly, leading to skills that activate too broadly and interfere with routing, reliability, and potentially security-sensitive task handling.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.