Back to skill

Security audit

OpenClaw Workspace Governance

Security checks for vulnerabilities and agentic risk

Overview

This skill is a disclosed workspace-governance playbook with local diagnostic helpers, with one hardening issue in a helper script but no evidence of hidden execution, exfiltration, persistence, or destructive behavior.

Reasonable to install if you want governance guidance for a complex OpenClaw workspace. Treat it as an operational playbook, not an auto-upgrader: review proposed process/routing changes before applying them. Only run freshness-check.py with trusted working-set configs, because a malicious config could make it inspect files outside the intended workspace.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/freshness-check.py:101
Finding

Workspace Root Escape Allows Out-of-Scope File Reads

Content
View full analysis

Vulnerability Details

File Location: scripts/freshness-check.py, lines 101 and 131–132
Vulnerability Type: Path traversal and insufficient filesystem-boundary validation
Risk Level: Medium

Complete Vulnerable Code Snippet

python
def extract_timestamp(path: Path) -> tuple[str, datetime, str]:
    text = path.read_text(encoding="utf-8", errors="ignore")
    if path.suffix == ".json":
        try:
            data = json.loads(text)
            if isinstance(data, dict) and "updatedAt" in data:
                dt = parse_dt(str(data["updatedAt"]))
                if dt:
                    return "updatedAt", dt, "json"
        except json.JSONDecodeError:
            pass
    head = "\n".join(text.splitlines()[:30])
    for pattern in DATE_PATTERNS:
        m = pattern.search(head)
        if m:
            dt = parse_dt(m.group(1))
            if dt:
                return "header-date", dt, "header"
    stat_dt = datetime.fromtimestamp(path.stat().st_mtime, tz=timezone.utc).astimezone()
    return "mtime", stat_dt, "filesystem"
python
def check_one(root: Path, thresholds: dict[str, int], item: FreshnessItem) -> FreshnessResult:
    if item.group not in thresholds:
        raise SystemExit(f"ERR: unknown group in working set: {item.group}")
    full_path = (root / item.path).resolve()
    if not full_path.exists():
        return FreshnessResult(
            group=item.group,
            path=item.path,
            thresholdDays=thresholds[item.group],
            source="missing",
            seenAt="",
            ageDays=999999.0,
            status="WARN",
            note="file missing",
        )
    label, seen_at, source = extract_timestamp(full_path)

Technical Analysis

The script accepts file paths from the configured workingSet and combines each path with the operator-provided workspace root. Although the resulting path is resolved, the code does not verify that the resolved target remains beneath that root ...[truncated 2752 chars]

Remediation
View remediation

Remediation Suggestions

Enforce a strict workspace containment boundary before accessing any configured target:

python
def resolve_workspace_file(root: Path, configured_path: str) -> Path:
    candidate = Path(configured_path)

    if candidate.is_absolute():
        raise ValueError(f"Absolute paths are not allowed: {configured_path}")

    resolved_root = root.resolve(strict=True)
    resolved_target = (resolved_root / candidate).resolve(strict=True)

    if not resolved_target.is_relative_to(resolved_root):
        raise ValueError(f"Path escapes workspace root: {configured_path}")

    if not resolved_target.is_file():
        raise ValueError(f"Target is not a regular file: {configured_path}")

    return resolved_target

Additional hardening should include:

  1. Reject absolute paths and traversal outside the resolved root.
  2. Apply containment checks after symlink resolution so an in-workspace symlink cannot target an external file.
  3. Require regular files and reject directories, devices, FIFOs, and other special objects.
  4. Apply a reasonable maximum file-size limit before reading.
  5. Read only the bounded prefix required for header-date extraction instead of loading the entire file.
  6. For JSON inputs, either enforce a conservative size limit before parsing or use a bounded data format appropriate for freshness metadata.
  7. Treat invalid or escaping paths as explicit configuration errors rather than silently processing them.
  8. Add regression tests covering:
    • ../ traversal
    • absolute paths
    • symlink escapes
    • valid nested workspace files
    • oversized files
    • special filesystem objects
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (13)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description substantially overstates and mischaracterizes the code. The actual script is narrowly scoped to scanning Markdown documentation for three recognized status labels and summarizing results. It is read-only, limited to local file inspection, and provides no evidence of system-level architecture changes, workspace upgrades, multi-agent governance, semantic search diagnostics, query routing, or maintenance automation. This is a material description-behavior mismatch, not merely an incomplete description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description presents this as a broad, high-impact workspace governance and upgrade skill with multiple advanced capabilities. The supplied code is much narrower: it only reads a config and target files, extracts timestamps, compares ages against thresholds, and outputs a report. While this loosely relates to 'working-set maintenance/freshness,' it does not perform governance, upgrades, routing, semantic diagnostics, or any system-level architecture changes. Therefore the description materially overstates and misrepresents the code's actual purpose and capabilities.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The description presents this as a high-risk, system-level workspace governance and upgrade skill for complex OpenClaw workspaces, implying substantial administrative or transformational behavior. The supplied code does not perform upgrades, governance, maintenance, routing, or multi-agent orchestration. Instead, it is a bounded diagnostic script for checking a semantic runtime by invoking an external command with status and query under supplied XDG paths and summarizing success/failure. The only partial overlap is 'semantic search diagnostics,' but even that is limited to basic runtime/status and representative query checks. Overall, the actual behavior is materially narrower and different from the declared primary purpose.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · README.md (reported line 2)May include surrounding context.

md
<p align="center">
  <img src="https://api.iconify.design/fluent-emoji:brain.svg" width="120" height="120" alt="Brain Icon" />
</p>

<h1 align="center">OpenClaw Workspace Governance</h1>

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
78% confidence
Finding

The document is primarily written in Chinese while only partially duplicating content in English, which can amount to an implicit language policy for users consuming the skill. The file does not state that language is user-selectable or that the Chinese-first presentation is required for a specific region or compliance context.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/query-class-routing.md (reported line 138)May include surrounding context.

md
Semantic search role:
- useful discovery layer
- often helpful for locating candidate files
- never enough by itself for definitive historical claims without verification

---

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The top-level description and warning are presented in Chinese-first bilingual form, and the rest of the document is largely duplicated with a Chinese section followed by an English section. While not overtly forbidding other languages, the skill does not explicitly offer a user language preference or opt-in, which can conflict with a language-choice policy if skills must not impose a locale by default.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The document presents headings and descriptive content in Chinese first and English second, which signals a fixed bilingual/localized format rather than a user-selectable language preference. Under the stated policy, language constraints should either provide user choice or be clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

This markdown reference mixes English content with several example trigger/query phrases that are only presented in Chinese, such as the system-state example on L035. Under the policy rule, language-specific phrasing can be a concern when the skill does not explicitly offer user language choice or explain that it is intentionally region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The governance examples include Chinese-only phrases alongside English ones, but the document does not state whether users may interact in any language or whether Chinese is required for certain routes. This creates a mild language-policy ambiguity in natural-language instructions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
74% confidence
Finding

At least some query-class examples, such as the runtime-now example on L073, are only shown in Chinese, and similar patterns recur in later sections. Because the file does not explicitly state that the skill supports multiple languages or that Chinese examples are illustrative only, this may imply an undocumented language preference.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
71% confidence
Finding

Examples like the user-preference phrase on L107 are presented only in Chinese, which may signal a default language expectation absent user opt-in. The policy allows locale constraints when clearly documented and justified, but this file does not provide that explanation.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.