Back to skill

Security audit

OpenClaw Harness

Security checks for vulnerabilities and agentic risk

Overview

This skill is not clearly malicious, but it advertises broad workspace-management abilities while shipping incomplete or unsafe tooling that can affect local state and, in one path, turn workspace configuration into shell execution.

Review this carefully before installing. Do not run the GC agent or examples on important workspaces until numeric config validation, symlink handling, path validation, restore warnings, and memory-file scoping are fixed. Treat .harness/config.json and any skill directory you package as trusted input only.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
bin/gc-agent.sh:96
Finding

Shell Command Injection Through Unvalidated GC Configuration

Content
View full analysis
/dev/null) echo "${val:-$default}" else echo "$default" fi } MAX_CP=$(get_gc_config "gc.max_checkpoints" 10) MAX_AGE=$(get_gc_config "gc.max_age_days" 7) REPORT_RETENTION=$(get_gc_config "gc.report_retention_count" 20) COMPRESS_TRIGGER=$(get_gc_config "gc.compress_trigger_lines" 200) TRASH_RETENTION=$(get_gc_config "gc.trash_retention_days" 30) ``` The resulting values are subsequently evaluated as Bash arithmetic expressions: ```bash local age_seconds=$((MAX_AGE * 86400)) ``` ```bash if [[ $cps_count -gt $MAX_CP ]]; then local excess=$(($cps_count - MAX_CP)) local to_delete=$(echo "$cps_list" | sort -t: -k2 | head -$excess) ``` ```bash local max_age_seconds=$((TRASH_RETENTION * 86400)) ``` ### Technical Analysis Numeric settings read from the workspace-controlled `.harness/config.json` file are not validated before they are inserted into Bash arithmetic contexts. Bash does not treat variable contents in arithmetic expansion strictly as decimal data. Instead, values may be recursively interpreted as arithmetic expressions. Crafted expressions, particularly expressions using array subscripts containing command substitutions, can therefore cause shell commands to execute when the GC process evaluates values such as `MAX_AGE`, `MAX_CP`, or `TRASH_RETENTION`. The affected code also uses unvalidated values in numeric comparisons and as the argument to `head`, increasing the likelihood of unexpected command behavior or denial of service even when a payload does not achieve comman ...[truncated 1450 chars]
Remediation
View remediation
max )); then log_error "$name is outside the permitted range" exit 1 fi } ``` 2. Apply validation immediately after configuration loading and before any arithmetic expansion: ```bash validate_uint "gc.max_checkpoints" "$MAX_CP" 0 10000 validate_uint "gc.max_age_days" "$MAX_AGE" 0 36500 validate_uint "gc.report_retention_count" "$REPORT_RETENTION" 0 100000 validate_uint "gc.compress_trigger_lines" "$COMPRESS_TRIGGER" 1 10000000 validate_uint "gc.trash_retention_days" "$TRASH_RETENTION" 0 36500 ``` 3. Reject invalid configuration rather than silently substituting or evaluating it. 4. Validate `DAEMON_INTERVAL` and all numeric command-line arguments using the same approach. 5. Use `10#$value` after validation to force decimal interpretation and avoid octal parsing. 6. Add tests containing whitespace, negative values, excessive values, arithmetic operators, array syntax, and command-substitution payloads. ]]>

T09 · Insecure Skill Coding Practices

Error
Location
scripts/package_skill.py:98
Finding

Symlink-Based Disclosure of Files Outside the Skill Directory

Content
View full analysis
/home/victim/.ssh/id_rsa ``` 3. The Skill contains a valid `SKILL.md`, so validation succeeds. 4. The victim runs: ```bash python scripts/package_skill.py /path/to/malicious-skill ``` 5. `rglob('*')` discovers the link, and `is_file()` reports that its target is a file. 6. `zipf.write()` follows the link and reads the external target. 7. The generated `.skill` archive contains the sensitive t ...[truncated 687 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/init_skill.py:205
Finding

Path Traversal Through Unvalidated Skill Names

Content
View full analysis
Remediation
View remediation
40: print("Error: Skill name exceeds 40 characters") return None ``` 2. Explicitly reject absolute paths and all path separators: ```python candidate_name = Path(skill_name) if candidate_name.is_absolute() or len(candidate_name.parts) != 1: print("Error: Skill name must not contain path components") return None ``` 3. Resolve the final destination and verify containment: ```python output_root = Path(path).resolve() skill_dir = (output_root / skill_name).resolve() if skill_dir.parent != output_root: print("Error: Destination escapes the selected output directory") return None ``` 4. Apply validation inside `init_skill()` rather than only in the CLI entry point, ensuring callers importing the function receive the same protection. 5. Create the output root separately and use exclusive directory creation for the final direct child. 6. Add tests for `..`, `../name`, absolute paths, repeated separators, Windows-style separators, empty names, leading or trailing hyphens, and symbolic-link output roots. 7. Remove partially created output if a later file-generation step fails, preventing incomplete executable trees from being left behind. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
Findings (28)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

If the real implementation only validates skill metadata while claiming checkpointing, restore, GC, progress tracking, and linting of agent files, users may invoke it under false assumptions and bypass appropriate review of what it actually does. Security controls depend on accurate disclosure; materially misleading capability statements create operational risk and can mask unexpected file-processing behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

If the real implementation only validates skill metadata while claiming checkpointing, restore, GC, progress tracking, and linting of agent files, users may invoke it under false assumptions and bypass appropriate review of what it actually does. Security controls depend on accurate disclosure; materially misleading capability statements create operational risk and can mask unexpected file-processing behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

If the real implementation only validates skill metadata while claiming checkpointing, restore, GC, progress tracking, and linting of agent files, users may invoke it under false assumptions and bypass appropriate review of what it actually does. Security controls depend on accurate disclosure; materially misleading capability statements create operational risk and can mask unexpected file-processing behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the real implementation only validates skill metadata while claiming checkpointing, restore, GC, progress tracking, and linting of agent files, users may invoke it under false assumptions and bypass appropriate review of what it actually does. Security controls depend on accurate disclosure; materially misleading capability statements create operational risk and can mask unexpected file-processing behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the real implementation only validates skill metadata while claiming checkpointing, restore, GC, progress tracking, and linting of agent files, users may invoke it under false assumptions and bypass appropriate review of what it actually does. Security controls depend on accurate disclosure; materially misleading capability statements create operational risk and can mask unexpected file-processing behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the real implementation only validates skill metadata while claiming checkpointing, restore, GC, progress tracking, and linting of agent files, users may invoke it under false assumptions and bypass appropriate review of what it actually does. Security controls depend on accurate disclosure; materially misleading capability statements create operational risk and can mask unexpected file-processing behavior.

Content

No source excerpt is available for this finding.

Chaining Abuse

High
Category
Tool Misuse
Confidence
75% confidence
Finding

Tool calls are chained to bypass individual safety checks or escalate capabilities beyond what any single tool call would allow.

Content

Scanner excerpt · examples/gc-agent-setup.sh (reported line 42)May include surrounding context.

sh
# 清理演示目录(幂等)
# -----------------------------------------------------------------------------
cleanup_demo() {
  [[ -d "$PROJECT" ]] && rm -rf "$PROJECT"
  info "演示目录已清理"
}

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The plugin design sources arbitrary local shell code from .harness/plugins/*/plugin.sh using Bash source, which executes with the full privileges of the current user. In an agent skill intended for cross-session automation, this creates a powerful code-execution extension point that can be abused by any untrusted plugin dropped into the workspace, enabling command execution, data exfiltration, persistence, or tampering.

Content

No source excerpt is available for this finding.

Self-Modification

High
Category
Rogue Agent
Confidence
85% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · scripts/init_skill.py (reported line 266)May include surrounding context.

python
# Print next steps
    print(f"\n✅ Skill '{skill_name}' initialized successfully at {skill_dir}")
    print("\nNext steps:")
    print("1. Edit SKILL.md to complete the TODO items and update the description")
    print("2. Customize or delete the example files in scripts/, references/, and assets/")
    print("3. Run the validator when ready to check the skill structure")

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill advertises commands that imply shell execution, file reads/writes, and potentially network use, but it does not declare any explicit tool scope such as permissions or allowed-tools. That makes the trust boundary unclear and can cause an agent platform to grant broader capabilities than reviewers expect, increasing the risk of unintended filesystem changes or command execution.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script performs a state-changing restore with --force and the interactive confirmation is explicitly commented out, so running the example can overwrite the current workspace without an active user decision at that moment. In a checkpoint-management skill, restore operations are expected, but making them unconditional in an example script increases the risk of accidental data loss or rollback of legitimate work if executed blindly or adapted into automation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This shell script contains natural-language descriptions, prompts, and operational guidance exclusively in Chinese, including the execution confirmation prompt. Under the language/locale policy, forcing a specific language without user opt-in is a policy violation unless the locale restriction is clearly documented and justified.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

Line L015 states the design is '无外部 API,纯本地执行' (no external API, purely local execution). Later, L509-L510 document a publish command to Clawhub, which introduces an external-facing action that contradicts that stated design principle.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This markdown file documents harness checkpoint restore <cp-id> with --force to overwrite current files, which can affect user data. Under the markdown variant of SQP-2, destructive or data-affecting behavior should be accompanied by an explicit warning, but the section lists the command semantics without a user-facing caution about overwrite risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The markdown specifies automatic cleanup and deletion behavior for old checkpoints, reports, and temporary files, including an aggressive mode. Although it mentions archival and logging, it does not clearly warn users that retained artifacts may be removed and that aggressive cleanup can affect recoverability or stored work products.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest describes checkpoint/snapshot management, verify/gc/restore, cross-session progress tracking, and linting agent configuration files. This design also introduces Git pre-commit hook installation and enforcement, which modifies repository workflow and is not an obvious requirement of the stated purpose.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The manifest emphasizes a local cross-session context manager and local linting/verification workflows. The documented harness publish --skill ... capability implies packaging and publishing to an external service, which is not part of the stated purpose and conflicts with the document’s earlier 'zero external dependency' design intent.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documented custom verification example allows arbitrary command execution via a user-supplied rule, which expands the tool from passive context management into an execution surface. In an agent setting, this is dangerous because untrusted configuration or prompts could cause the harness to run attacker-chosen local commands during verify or through automated hooks.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The documentation presents rm -rf .harness as a repair step that irreversibly deletes stored checkpoints and state. Although it is labeled as dangerous, the guidance is still easy to copy-paste, and in this skill's context the .harness directory contains the very cross-session memory and recovery data the tool is meant to protect, making accidental destructive loss more likely.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · references/maintenance.md (reported line 85)May include surrounding context.

修复:

bash
# Debian/Ubuntu
sudo apt install jq

# macOS
brew install jq

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The requirements specify checkpoint restore capability but do not explicitly warn that restore can overwrite the current workspace state, replacing files with older snapshot contents. In this skill's context, restore is especially dangerous because it manages cross-session progress and task artifacts; a mistaken restore could silently discard newer work, corrupt task state, or reintroduce stale configuration across an agent workflow.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The requirements describe automatic cleanup and deletion behavior for checkpoints, reports, and temporary files, but they do not explicitly require interactive confirmation, prominent user-facing risk warnings, or clear safeguards before destructive actions beyond dry-run support. In a context-management skill that operates across sessions and maintains task state, omission of explicit destructive-operation warnings increases the chance of unintended data loss or premature cleanup, especially if users run GC with non-default retention settings.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest describes a harness for checkpoints, verification, garbage collection, restore, progress tracking, and linting agent configuration files. The 'Scripts' section additionally documents helper tools for initializing entirely new skills and packaging arbitrary skills, which is a broader developer-tooling function not implied by the manifest description.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

All user-facing comments and terminal messages are written in Chinese, with no indication that the skill is region-specific or that users can choose another language. This creates a language/locale constraint in the skill's natural-language interface without opt-in or justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

SQP-3 applies to all file types and covers language or locale policy violations. This architecture document consistently uses Chinese for headings, instructions, CLI behavior descriptions, and conclusions, but does not indicate that the language choice is optional or justified as region-specific.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.