Back to skill

Security audit

Ride Receipts

Security checks for vulnerabilities and agentic risk

Overview

This skill’s purpose is coherent, but it handles highly sensitive Gmail receipt data and Gateway tokens with some under-protected data paths that warrant review before installation.

Review before installing if your ride receipts are sensitive. Prefer a loopback Gateway URL, avoid private or remote plaintext HTTP Gateway targets, use a short-lived token if available, store outputs in a private directory, and be cautious opening or sharing the generated CSV in spreadsheet software.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/extract_rides_gateway.py:67
Finding

Raw receipt emails and Gateway bearer token may be transmitted over plaintext HTTP

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/extract_rides_gateway.py:103
Finding

Untrusted email content is concatenated with extraction instructions without a trust boundary

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/export_anonymized_rides_csv.py:127
Finding

The anonymized CSV exporter does not neutralize spreadsheet formulas

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/fetch_emails_json.py:108
Finding

Sensitive ride artifacts are created without restrictive file permissions

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (9)

Tainted flow: 'req' from os.environ.get (line 109, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
93% confidence
Finding

The request sent via urllib.request.urlopen is built from environment/config-controlled values, including OPENCLAW_GATEWAY_URL and the bearer token, and the code transmits full ride-receipt email contents to that endpoint. Although the script attempts to restrict hosts to local/private addresses by default, that protection can be bypassed with OPENCLAW_ALLOW_NONLOCAL_GATEWAY=1, so a malicious or misconfigured environment can exfiltrate sensitive email data and credentials to an attacker-controlled service.

Content

Scanner excerpt · scripts/extract_rides_gateway.py (reported line 118)May include surrounding context.

python
},
        method='POST',
    )
    with urllib.request.urlopen(req, timeout=timeout) as resp:
        body = json.loads(resp.read().decode('utf-8'))
    for item in body.get('output', []):
        if item.get('type') == 'message':

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is for an ingestion pipeline: fetch ride receipt emails from Gmail, use a Gateway-backed model to extract ride details, save intermediate JSON files, and insert data into SQLite. The supplied code does none of that. Instead, it assumes an already-populated SQLite database exists, queries the rides table, derives anonymized/normalized fields such as email month and 15-minute-rounded times, optionally parses metrics from extracted_ride_json, and outputs a CSV. This is a materially different purpose and capability set from the declared behavior, so it is a clear mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
87% confidence
Finding

The implemented script is narrowly an extraction utility. It accepts --emails-json and --out, loads preexisting email data, calls a Gateway endpoint with each email, parses returned JSON, and writes the accumulated ride records to an output JSON file. It includes retry/resume behavior and host restrictions for the Gateway, but no Gmail access, no gog integration, and no SQLite operations. Because the declared description presents a broader pipeline whose key stages include email fetching and SQLite insertion, the description materially overstates what this code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents a complete pipeline: fetch Gmail ride receipts, run Gateway-backed extraction over each email, produce rides.json, and load the results into SQLite. The supplied code chunk is much narrower. It searches Gmail using gog, fetches message metadata and raw bodies, parses MIME content, collects HTML, and writes the resulting email records to a JSON file. There is no Gateway/OpenClaw interaction, no receipt-to-ride extraction logic, no rides.json output, and no SQLite usage. While the fetch behavior is consistent with one component of the description, the actual code does not match the declared end-to-end purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description presents a complete pipeline for fetching ride receipt emails, extracting structured ride data with a Gateway-backed LLM, and loading that data into SQLite. The supplied code chunk does not implement any of those email-fetching, LLM-extraction, JSON-processing, or ride-ingestion behaviors. It only initializes a SQLite database from a schema file. While database setup could be a supporting component of the broader system, the chunk’s actual purpose is materially narrower and different from the declared end-to-end pipeline, so this is a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description claims an end-to-end ingestion pipeline from Gmail to SQLite using gog and a Gateway LLM extraction step. The actual code does none of those things. It accepts an existing SQLite database path, reads the rides table schema and contents, runs aggregate SQL queries, and outputs a JSON summary. This is a materially different primary purpose and omits the core declared behaviors (email fetching, LLM extraction, JSON artifact generation, and database population), so this is a clear description-behavior mismatch.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill declares capabilities that include shell, network, filesystem, and environment access, but it does not define an explicit tool scope such as permissions or allowed-tools. That creates an overbroad execution surface where an agent may use more power than the user expects, especially because the workflow handles sensitive Gmail receipt data and Gateway auth tokens.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The manifest describes a pipeline that fetches ride-receipt emails, extracts ride data with a Gateway-backed LLM, writes rides.json, and inserts results into SQLite. This script adds a separate data-export capability that reads the ride database and emits a CSV containing provider, fare, timing, and pickup/dropoff location fields, which is not justified by that stated purpose.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/fetch_emails_json.py (reported line 18)May include surrounding context.

python
def run_json(cmd):
    out = subprocess.check_output(cmd, text=True)
    return json.loads(out)

Static analysis

No suspicious patterns detected.