Back to skill

Security audit

Claw Mechanic

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent OpenClaw troubleshooting skill that can touch sensitive host configuration, but its risky actions are disclosed, scoped to repair work, and gated with approval/redaction guidance.

Install this only for agents you trust to operate an OpenClaw host. It is appropriate for diagnostics and repair, but before allowing changes confirm the target host, permitted change scope, backup location, and whether restarts, plugin updates, cron edits, or credential-adjacent checks are allowed. Do not permit raw secret, env-file, session-body, or private workspace inspection unless you explicitly approve that exact step.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
Findings (9)

YARA rule 'agent_skill_prompt_injection_hidden_instructions': Prompt injection or hidden instructions embedded in AI agent skill text [agent_skills]

High
Category
YARA Match
Confidence
80% confidence
Finding

YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

Content

Scanner excerpt · references/failure-map.md (reported line 11)May include surrounding context.

md
# Claw Mechanic Failure Map

Use this reference when the quick workflow points at a specific subsystem. Keep output evidence short and exact.

## Docs To Check

- CLI overview: `https://docs.openclaw.ai/cli`
- Doctor: `https://docs.openclaw.ai/doctor`
- Plugins: `https://docs.openclaw.ai/cli/plugins`
- Cron: `https://docs.openclaw.ai/cli/cron`
- Models: `https://docs.openclaw.ai/cli/models`
- Secrets: `https://docs.openclaw.ai/cli/secrets`
- Exec approvals: `https://docs.openclaw.ai/tools/exec-approvals`
- Gateway APIs/config: `https://docs.openclaw.ai/gateway/configuration`
- Channel behavior: check the relevant OpenClaw channel docs under `https://docs.openclaw.ai/channels/`

Docs source order:
- Record `openclaw --version` first so version drift is visible.
- Prefer installed package docs when exact installed behavior matters; use current `https://docs.openclaw.ai` for live docs and generated/API pages.
- If using a downloaded `llm

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
90% confidence
Finding

Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.

Content

Scanner excerpt · references/failure-map.md (reported line 205)May include surrounding context.

md
Fix pattern:
- Verify all agents, not just `main`.
- Confirm provider, model, vector dimensions, semantic availability, and FTS.
- For context-engine plugins, ensure summary/expansion models are allowed by plugin `llm`/`subagent` override policy.
- If generated low-value sessions are noisy, use the plugin's documented ignore/stateless session patterns when supported, and keep productive channel sessions enabled.
- If plugin-owned state is already polluted, back up the plugin database/state first, prefer archiving or retiring generated active rows over deleting raw history, then restart and verify no new warnings.
- `openclaw sessions cleanup` repairs OpenClaw session stores/transcripts; do not assume it cleans plugin-owned context databases.

Self-Modification

High
Category
Rogue Agent
Confidence
90% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · references/root-failure-taxonomy.md (reported line 195)May include surrounding context.

md
Mechanic move: keep noncritical broken backends out of active routes until fixed. Do not churn OpenClaw config when the upstream is the proven fault.

## 13. Self-Update And Restart Hazard

Root failure: an agent updated the host or plugins while depending on the gateway it was changing, leaving installed package state and running service state out of sync.

Self-Modification

High
Category
Rogue Agent
Confidence
90% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · references/root-failure-taxonomy.md (reported line 227)May include surrounding context.

md
- If access, service ownership, or lifecycle is unproven, fix that first.
- If one layer is green but user-visible behavior is bad, test runtime, cron, tasks, channels, and models separately.
- If cost is the symptom, prioritize cron payload models, fallbacks, context load, reviewer model, and retry/timeout loops.
- If the host just updated, prioritize plugin cohort, migration drift, service lifecycle, and self-update hazards before model rewrites.
- If local model APIs are involved, prove the exact OpenClaw runtime path, not only a generic HTTP request.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The README says to invoke the skill when a host is 'slow, looping, costly, stale, broken after update, or behaving differently,' which are broad conditions that can overlap with many ordinary troubleshooting contexts. It does not provide explicit boundaries or negative examples clarifying when this skill should not be invoked, increasing the risk of unintended activation.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 25)May include surrounding context.

md
- Before docs-sensitive config changes, record `openclaw --version`; prefer installed package docs when you need exact-version behavior, then current official docs at `https://docs.openclaw.ai`. Use docs for command semantics, plugin behavior, cron model/fallback behavior, approvals, secrets, and gateway APIs.
- Do not trust one green command. Pair static checks with live proof: gateway health, plugin doctor, task/cron state, model smoke, channel status, or direct endpoint probes as appropriate.
- Back up config/state before edits. Prefer `openclaw config set`, Gateway `config.patch`, or other supported CLI/Gateway RPC paths over direct file/database edits. Use direct DB/file edits only when the live API has no path, and verify by re-reading through OpenClaw afterward.
- Keep repairs small. Do not update OpenClaw, reinstall plugins, restart gateways, or rewrite model routes unless the evidence points there or the user approved that class of change.
- Separate confirmed findings from stale/historical warnings and from audit-tool false positives.
- Do not read raw secrets, env files, session bodies, or agent workspace contents unless the user approves that exact inspection. Prefer audits, redacted probes, metadata, and config paths over secret values.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · references/failure-map.md (reported line 63)May include surrounding context.

md
- Service-manager restart counts and recent stdout/stderr logs.

Fix pattern:
- Do not batch-rewrite `openclaw.json`. Use `openclaw config set` for single keys or Gateway `config.patch` with a fresh `baseHash` for structured edits.
- Back up the exact config source file first. If `$include` files are involved, edit the source include file, not a flattened copy.
- JSON5 syntax, env references, and `$include` layout can be valid config features; validate before removing them.
- Most channels, agents, models, cron, tools, bindings, session, messages, and logging changes hot-apply. `gateway.*`, `discovery`, and `plugins` usually need restart or hybrid restart handling; verify against current docs before restarting.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The guidance explicitly tells operators to inspect service environment files such as ~/.openclaw/.env, service-env files, and launchctl environment, which commonly contain API keys, channel tokens, and other credentials. Although it says not to print secret values in one narrow case, it does not consistently require redaction or least-exposure handling before reviewing these files, so an agent following the playbook could expose secrets in logs, transcripts, or user-visible output.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

The instruction to 'Keep output evidence short and exact' is neutral, but the document as written is entirely prescriptive in English and provides no option for locale or language choice. Under the stated policy, skills should not force a specific language without user opt-in or a documented justification.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.