Back to skill

Security audit

Openclaw Warden

Security checks for vulnerabilities and agentic risk

Overview

This is a local workspace security scanner, but it includes under-disclosed commands that can overwrite important workspace files or disable installed skills without confirmation.

Install only if you are comfortable with a local tool that reads your agent workspace and writes .integrity baseline data. Avoid using protect, restore, rollback, or quarantine unless you have reviewed the exact target paths and have backups. Create baselines only on a known-clean workspace, and treat clean scan results cautiously because some security-skill names are exempted by directory name.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/integrity.py:625
Finding

Workspace Boundary Traversal in File Acceptance and Snapshot Restoration

Content
View full analysis
Path | None: """Get the snapshot path for a file, or None if no snapshot exists.""" p = snapshot_dir(workspace) / rel return p if p.is_file() else None ``` ### Technical Analysis The user-provided file path is only normalized by replacing backslashes with forward slashes. The implementation does not reject absolute paths, `..` components, symlinks, or resolved paths outside the configured workspace. In `cmd_accept`, joining an absolute path to `workspace` can discard the workspace prefix under `pathlib` semantics. A path containing traversal components can similarly resolve outside the workspace. The command can therefore hash and register arbitrary readable files rather than only files belonging to the monitored workspace. The rest ...[truncated 2276 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/integrity.py:525
Finding

Name-Based Security Scanner Exemption Allows Injection Detection Bypass

Content
View full analysis
list[dict]: """Scan all workspace files for injection patterns.""" files = collect_monitored_files(workspace) all_findings = [] for rel, abspath in sorted(files.items()): # Skip security skills that document injection patterns — # their SKILL.md will always trigger false positives if rel.startswith("skills/"): parts = rel.split("/") if len(parts) >= 2 and parts[1] in SECURITY_SCAN_EXEMPT: continue ``` The same exemption also prevents automated quarantine: ```python if rel.startswith("skills/") and "/SKILL.md" in rel: # Extract skill name: skills//SKILL.md parts = rel.split("/") if len(parts) >= 2: skill_name = parts[1] if skill_name.startswith(QUARANTINE_PREFIX): continue if skill_name in SECURITY_SCAN_EXEMPT: continue ``` ### Technical Analysis The scanner treats a skill as trusted solely because its directory name appears in `SECURITY_SCAN_EXEMPT`. It does not verify package provenance, a cryptographic signature, a known content hash, repository identity, or an administrator-controlled trust record. Directory names inside the workspace are not reliable security identities. An attacker who can install, replace, or rename a skill can use `openclaw-warden` or `openclaw-bastion` and receive a complete exempt ...[truncated 1820 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
`. 8. The legitimate changes are lost. If the baseline was poisoned, attacker-selected historical content is restored. ### Impact Assessment The command can overwrite all modified critical workspace files using the invoking user's filesystem permissions. This can cause: - Loss of legitimate identity, policy, heartbeat, or tool configuration changes. - Restoration of stale or compromised content. - Unexpected behavioral changes in future Agent sessions. - Operational disruption caused by silent rollback. ...[truncated 411 chars]:781
Finding

Automated Protection Reverts All Modified Critical Files Without Confirmation

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (19)

YARA rule 'agent_skill_prompt_injection_hidden_instructions': Prompt injection or hidden instructions embedded in AI agent skill text [agent_skills]

High
Category
YARA Match
Confidence
80% confidence
Finding

YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

Content

Scanner excerpt · README.md (reported line 10)May include surrounding context.

nClaw Warden

Free workspace integrity verification for OpenClaw, Claude Code, and any Agent Skills-compatible tool.

Detects unauthorized modifications to agent identity and memory files and scans for prompt injection patterns — the post-installation security layer that other tools miss.

The Problem

AI agents read workspace files (SOUL.md, AGENTS.md, IDENTITY.md, memory files) on every session startup and trust them implicitly. Existing security tools scan skills before installation. Nothing monitors the workspace itself afterward.

A compromised skill, a malicious payload, or any process with file access can inject hidden instructions, embed exfiltration URLs, override safety boundaries, or plant persistent backdoors.

This skill detects all of these.

Install

bash
# Clone
git clone https://github.com/AtlasPA/openclaw-warden.git

# Copy to your workspace skills direc

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
90% confidence
Finding

Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.

Content

Scanner excerpt · README.md (reported line 12)May include surrounding context.

md
AI agents read workspace files (`SOUL.md`, `AGENTS.md`, `IDENTITY.md`, memory files) on every session startup and **trust them implicitly**. Existing security tools scan *skills* before installation. Nothing monitors the *workspace itself* afterward.

A compromised skill, a malicious payload, or any process with file access can inject hidden instructions, embed exfiltration URLs, override safety boundaries, or plant persistent backdoors.

This skill detects all of these.

Instruction Override

High
Category
Prompt Injection
Confidence
90% confidence
Finding

This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

Content

Scanner excerpt · README.md (reported line 12)May include surrounding context.

md
AI agents read workspace files (`SOUL.md`, `AGENTS.md`, `IDENTITY.md`, memory files) on every session startup and **trust them implicitly**. Existing security tools scan *skills* before installation. Nothing monitors the *workspace itself* afterward.

A compromised skill, a malicious payload, or any process with file access can inject hidden instructions, embed exfiltration URLs, override safety boundaries, or plant persistent backdoors.

This skill detects all of these.

Instruction Override

High
Category
Prompt Injection
Confidence
80% confidence
Finding

This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

Content

Scanner excerpt · README.md (reported line 52)May include surrounding context.

md
- New untracked files

### Prompt Injection Patterns
- **Instruction override** — "ignore previous instructions", "you are now", "forget your instructions"
- **System prompt markers** — `<system>`, `[SYSTEM]`, `<<SYS>>`, `[INST]`
- **Markdown exfiltration** — Image tags with encoded data in URLs
- **Base64 payloads** — Large encoded blobs outside code blocks

Instruction Override

High
Category
Prompt Injection
Confidence
80% confidence
Finding

This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 93)May include surrounding context.

md
- New untracked files

### Prompt Injection Patterns
- **Instruction override** — "ignore previous instructions", "you are now", "forget your instructions"
- **System prompt markers** — `<system>`, `[SYSTEM]`, `<<SYS>>`, `[INST]`
- **Markdown exfiltration** — Image tags with encoded data in URLs
- **Base64 payloads** — Large encoded blobs outside code blocks

Instruction Override

High
Category
Prompt Injection
Confidence
80% confidence
Finding

This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

Content

Scanner excerpt · README.md (reported line 52)May include surrounding context.

md
- New untracked files

### Prompt Injection Patterns
- **Instruction override** — "ignore previous instructions", "you are now", "forget your instructions"
- **System prompt markers** — `<system>`, `[SYSTEM]`, `<<SYS>>`, `[INST]`
- **Markdown exfiltration** — Image tags with encoded data in URLs
- **Base64 payloads** — Large encoded blobs outside code blocks

YARA rule 'agent_skill_prompt_injection_hidden_instructions': Prompt injection or hidden instructions embedded in AI agent skill text [agent_skills]

High
Category
YARA Match
Confidence
80% confidence
Finding

YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

Content

Scanner excerpt · SKILL.md (reported line 43)May include surrounding context.

ty

Check all monitored files against the stored baseline. Reports modifications, deletions, and new untracked files.

bash
python3 {baseDir}/scripts/integrity.py verify --workspace /path/to/workspace

Scan for Injections

Scan workspace files for prompt injection patterns: hidden instructions, base64 payloads, Unicode tricks, markdown image exfiltration, HTML injection, and suspicious system prompt markers.

bash
python3 {baseDir}/scripts/integrity.py scan --workspace /path/to/workspace

Full Check (Verify + Scan)

Run both integrity verification and injection scanning in one pass.

bash
python3 {baseDir}/scripts/integrity.py full --workspace /path/to/workspace

Quick Status

One-line summary of workspace health.

bash
python3 {baseDir}/scripts/integrity.py status --workspace /path/to/workspace

Accept Changes

After reviewing a legitimate change, update the baseline for a specific file.

bash
python3 {baseDir}/scripts/integrity.py accept

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

Protect mode chains multiple destructive actions—restore, git rollback, and skill quarantine—based only on scan heuristics and without user confirmation. In this skill context, that is especially dangerous because the scanner intentionally searches for prompt-injection phrases; an attacker can plant trigger text to induce automated changes, causing denial of service or integrity loss in the workspace.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The README instructs users to run baseline and accept operations that necessarily update the trusted integrity state, but it does not warn that doing so can legitimize already-compromised files if performed at the wrong time. In an integrity tool, silent trust-state updates are security-sensitive because they can convert an active compromise into the new accepted baseline and suppress future alerts.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill advertises executable commands that invoke a local Python script with workspace access, but the manifest does not declare any tool scope such as permissions or allowed-tools. This creates a capability/visibility mismatch: users or hosting agents may not realize the skill can read environment variables, traverse workspace files, and update integrity state, increasing the chance of overbroad execution and unsafe trust.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 27)May include surrounding context.

Establish Baseline

Create or reset the integrity baseline. Run this after setting up your workspace or after reviewing and accepting all current file states.

bash
python3 {baseDir}/scripts/integrity.py baseline --workspace /path/to/workspace

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The restore command overwrites workspace files directly from snapshots without a confirmation prompt, backup, or pre-restore diff. If invoked on the wrong file or under adversarial prompting, it can silently destroy legitimate updates and revert trusted content to stale baseline state.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill can execute git-based rollback operations against workspace files, which gives it the ability to alter user content based on runtime inputs. Even though the git commands are not shell-injectable, this creates a dangerous integrity boundary: a mistaken or manipulated invocation can discard legitimate local changes or restore attacker-chosen repository state if the repository itself is untrusted.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The rollback command performs a destructive git checkout of a specified file with no confirmation and no preview of the discarded changes. In a tool intended to react to suspicious content, this can be abused socially or accidentally to wipe user work or force reversion to an untrusted repository state.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/integrity.py (reported line 667)May include surrounding context.

python
sys.exit(1)

    # Check if file is tracked by git
    result = subprocess.run(
        ["git", "ls-files", rel],
        cwd=str(workspace),
        capture_output=True, text=True,

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/integrity.py (reported line 677)May include surrounding context.

python
sys.exit(1)

    # Checkout from HEAD
    result = subprocess.run(
        ["git", "checkout", "HEAD", "--", rel],
        cwd=str(workspace),
        capture_output=True, text=True,

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Protect mode automatically restores files from snapshots, performs git checkout, and renames skill directories as a countermeasure. Because these actions are triggered by heuristic scan results and happen without approval, a false positive or attacker-crafted content can cause unauthorized workspace modification or denial of service by quarantining legitimate skills.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/integrity.py (reported line 795)May include surrounding context.

python
git_dir = workspace / ".git"
            if git_dir.exists():
                import subprocess
                result = subprocess.run(
                    ["git", "checkout", "HEAD", "--", rel],
                    cwd=str(workspace),
                    capture_output=True, text=True,

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

Quarantine renames an entire skill directory immediately, disabling it on the next session without advance confirmation. This is less severe than file overwrite, but it can still be used to disrupt legitimate functionality or create avoidable operational outages if triggered accidentally or through misleading guidance.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.prompt_injection_instructions

Prompt-injection style instruction pattern detected.

Warn
Code
suspicious.prompt_injection_instructions
Location
README.md:52

Prompt-injection style instruction pattern detected.

Warn
Code
suspicious.prompt_injection_instructions
Location
SKILL.md:93