Back to skill

Security audit

nutcrackertest

Security checks for vulnerabilities and agentic risk

Overview

This skill is purpose-aligned, but it reads and stores local conversation history and has incomplete redaction and date-scoping safeguards.

Install only if you are comfortable with the skill reading your local OpenClaw conversation transcripts and writing redacted session data, reports, quotes, behavioral summaries, and trend files on disk. Avoid using it on histories that may contain secrets or sensitive personal data until the date-filtering and path-redaction issues are fixed, and review/delete generated files in the skill's data and reports directories as needed.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/collect.sh:108
Finding

Date Filtering Occurs After Sensitive Session Metadata and Tool Content Are Collected

Content
View full analysis

Vulnerability Details

File Location: scripts/collect.sh:41-42, 108-145
Vulnerability Type: Insufficient data minimization and incomplete date filtering
Risk Level: Medium

Vulnerable Code

python
target_date = sys.argv[1]
files = sys.argv[2:]
python
# Count by role
session_data["message_count"] += 1
if role == "user":
    session_data["user_message_count"] += 1
elif role == "assistant":
    session_data["assistant_message_count"] += 1
elif role == "toolResult":
    session_data["tool_result_count"] += 1
    # Record tool call info
    tool_info = {
        "timestamp": timestamp,
        "content_preview": text[:200] if text else "",
        "success": "error" not in text.lower()[:500] if text else True
    }
    session_data["tool_calls"].append(tool_info)

# Store message
message_obj = {
    "role": role,
    "text": text,
    "timestamp": timestamp,
    "cost": cost
}
session_data["messages"].append(message_obj)

...

# Filter messages to only those on target date
session_data["messages"] = [
    m for m in session_data["messages"]
    if m.get("timestamp", "")[:10] == target_date
]

Technical Analysis

The collector reads and processes every message in every discovered session file before applying the requested date filter. Although the messages array is filtered at the end, the following fields are populated from the complete transcript and are not subsequently filtered or recalculated:

  • tool_calls, including up to 200 characters of tool-result content
  • message_count and role-specific counters
  • total_cost
  • session_start and session_end
  • Calculated session duration

Consequently, a dataset described as representing one date can retain conversation-derived content and metadata from other dates. The redaction stage may reduce some recognizable PII, but it does not eliminate the underlying unnecessary collection ...[truncated 1595 chars]

Remediation
View remediation

Remediation Suggestions

Apply the target-date check immediately after parsing each entry and before extracting message content or updating any aggregate:

python
timestamp = entry.get("timestamp", "")
if not timestamp or timestamp[:10] != target_date:
    continue

Then calculate messages, tool_calls, counters, cost, start/end timestamps, and duration exclusively from accepted entries. Additional hardening should include:

  1. Recalculate all aggregates after filtering rather than preserving whole-session values.
  2. Filter tool_calls by date even if early filtering is introduced, providing defense in depth.
  3. Avoid retaining tool-result previews unless they are strictly required for analysis.
  4. Add tests using a multi-day JSONL fixture and assert that no content, costs, timestamps, or counts from other dates appear in output.
  5. Define whether date matching uses UTC or local time and parse timestamps rather than relying solely on a ten-character prefix.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/redact.py:134
Finding

Absolute Session File Paths Bypass the PII Redaction Pipeline

Content
View full analysis

Vulnerability Details

File Location: scripts/collect.sh:44-45; scripts/redact.py:134-149
Vulnerability Type: Incomplete redaction of path-bearing fields
Risk Level: Medium

Vulnerable Code

The collector places the original session path into its output:

python
for filepath in files:
    session_data = {
        "file": filepath,
        "messages": [],
        "tool_calls": [],
        "session_start": None,
        "session_end": None,
        "total_cost": 0.0,
        "message_count": 0,
        "user_message_count": 0,
        "assistant_message_count": 0,
        "tool_result_count": 0,
        "has_target_date_messages": False
    }

The recursive redactor only sanitizes selected keys and passes file through unchanged:

python
def redact_session_data(data):
    """Recursively redact PII in session data structures."""
    if isinstance(data, str):
        return redact_text(data)
    elif isinstance(data, list):
        return [redact_session_data(item) for item in data]
    elif isinstance(data, dict):
        result = {}
        for key, value in data.items():
            if key in ("text", "content_preview"):
                result[key] = redact_text(value) if isinstance(value, str) else value
            elif key == "messages":
                result[key] = redact_session_data(value)
            elif key == "tool_calls":
                result[key] = redact_session_data(value)
            else:
                result[key] = value
        return result
    else:
        return data

Technical Analysis

collect.sh receives session filenames produced by find and stores each filename under the file property. When an absolute sessions directory is supplied, this property is an absolute path and can include a local username, agent identifier, directory layout, or other identifying filesystem information.

Although `re ...[truncated 1792 chars]

Remediation
View remediation

Remediation Suggestions

Do not persist the source path unless it is essential. Replace it with a non-sensitive identifier, basename, or stable hash. For example:

python
session_data = {
    "session_id": os.path.basename(filepath),
    ...
}

Strengthen the redactor so every string is recursively processed rather than relying on a small key allowlist:

python
elif isinstance(data, dict):
    return {
        key: redact_session_data(value)
        for key, value in data.items()
    }

Because generic string redaction may not recognize every sensitive path format, also apply explicit controls:

  1. Remove file before persistence, or replace its value with [PATH].
  2. Explicitly sanitize file, session_file, path, and error-message fields.
  3. Avoid including source filenames in friction, delight, and quote records.
  4. Add tests for Unix, macOS, and Windows home-directory paths.
  5. Test nested dictionaries and newly introduced fields to ensure future schema changes cannot bypass redaction.
  6. Restrict generated data and report files to owner-only permissions where supported.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Rogue AgentSelf-Modification, Session Persistence
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (6)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The skill's declared behavior does not align with what is actually specified, especially around access to local session files despite empty permissions. Such mismatches are dangerous because users may invoke the skill under false assumptions about what it does, while the implementation path still directs the agent to process sensitive conversation history and persist derived data.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill's declared behavior does not align with what is actually specified, especially around access to local session files despite empty permissions. Such mismatches are dangerous because users may invoke the skill under false assumptions about what it does, while the implementation path still directs the agent to process sensitive conversation history and persist derived data.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill instructs the agent to read local session transcripts, write analysis artifacts and reports, and execute shell/Python commands, but it declares no tool scope or permissions. This creates a transparency and containment failure: users and enforcement layers cannot accurately understand or restrict the skill's access to sensitive local data, increasing the risk of overbroad file access and silent data collection.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill is presented as a passive observer of OpenClaw usage and stores derived session data and reports, yet it lacks an upfront warning in the main description that local interaction history will be inspected and persisted. This undermines informed consent for a privacy-sensitive workflow and may surprise users into exposing historical conversations, including sensitive content, even if redaction is attempted.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
80% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · scripts/analyze.py (reported line 642)May include surrounding context.

python
if "automation" in tasks and tasks["automation"] > 0.2:
        recs.append({
            "priority": "medium",
            "recommendation": "Automation tasks are frequent. Explore OpenClaw's cron job feature to schedule recurring tasks automatically.",
        })

    # Archetype-based recommendations

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The template explicitly includes sections for verbatim quotes and behavioral archetyping/trends, which can encourage collection and reporting of personally sensitive behavioral data without any visible notice, minimization guidance, or consent boundary. In the context of a passive observation skill, this increases the risk of capturing identifiable user content, profiling users, and retaining sensitive interaction data in reports that may be shared more broadly than intended.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.