Back to skill

Security audit

Capability Evolver Pro.Bak

Security checks for vulnerabilities and agentic risk

Overview

This local log-analysis skill is purpose-aligned and does not show evidence of network exfiltration, persistence, credential access, or destructive behavior.

Install only if you are comfortable giving the skill the logs you pass to it. Redact secrets, tokens, personal data, and sensitive internal details first, and avoid wiring it to unbounded production or fleet log streams without request-size, runtime, and rate limits.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
handler.ts:94
Finding

Unenforced Log Input Limit Enables Resource Exhaustion

Content
View full analysis

Vulnerability Details

File Location: handler.ts:94-97 (validation flaw), with the unenforced limit declared at handler.ts:292-294
Vulnerability Type: Unbounded input processing / denial of service
Risk Level: Medium

Vulnerable Code

ts
if ((input.action === 'analyze' || input.action === 'evolve') &&
    (!input.logs || !Array.isArray(input.logs) || input.logs.length === 0)) {
  errors.push('"logs" array is required for analyze/evolve actions');
}

The status response advertises a limit that is not enforced:

ts
limits: {
  max_logs_per_request: 10000,
  max_patterns_returned: 50,
  max_recommendations: 20,
},

Technical Analysis

Request validation only verifies that logs is a non-empty array. It does not reject arrays exceeding the declared 10,000-entry maximum, limit the byte size of the request, constrain individual string fields, or validate each log entry's structure.

The accepted array is subsequently processed through multiple full-array scans, filtering operations, maps, sets, sorting, and string generation. The evolve action invokes the same analysis and performs additional processing. Consequently, an attacker-controlled oversized array—or entries containing extremely large message, context, timestamp, or stack strings—can cause disproportionate CPU and memory consumption.

This violates the implementation's own documented limit and leaves availability dependent on limits imposed by the surrounding runtime.

Attack Path

  1. An attacker or untrusted caller invokes the Skill using the analyze or evolve action.
  2. The attacker supplies a non-empty logs array containing substantially more than 10,000 entries, or supplies entries with very large string fields.
  3. validateRequest accepts the payload because it imposes no upper bound.
  4. handleAnalyze repeatedly traverses and aggregates the attacker-controlled data; evolve adds ...[truncated 703 chars]
Remediation
View remediation

Remediation Suggestions

  1. Enforce the declared maximum during validation:

    ts
    const MAX_LOGS = 10_000;
    
    if (
      (input.action === 'analyze' || input.action === 'evolve') &&
      (!Array.isArray(input.logs) ||
        input.logs.length === 0 ||
        input.logs.length > MAX_LOGS)
    ) {
      errors.push(`"logs" must contain between 1 and ${MAX_LOGS} entries`);
    }
    
  2. Validate every log entry, including allowed severity values, timestamp format, and required string fields.

  3. Apply explicit length limits to message, context, timestamp, and stack. Reject the request rather than silently accepting exceptionally large values.

  4. Enforce a maximum serialized request-body size at the OpenClaw or HTTP host boundary before parsing the full payload.

  5. Configure runtime memory, execution-time, concurrency, and per-caller rate limits to contain resource-exhaustion attempts.

  6. Add tests confirming rejection of arrays above 10,000 entries, oversized strings, malformed entries, sparse arrays, and payloads near the configured byte-size boundary.

Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (4)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The top-level description says to use the skill when the user or agent asks to analyze logs, diagnose failures, improve reliability, generate evolution proposals, or assess system health. These are broad operational phrases that can match ordinary conversations, making unintended invocation more likely. Because the skill is designed to process runtime logs, accidental triggering can pull sensitive production telemetry into the skill flow without a clear, explicit opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill heavily promotes analyzing production, staging, and fleet-wide logs, but it does not warn that logs often contain secrets, tokens, personal data, internal paths, or other sensitive telemetry. Even if processing is local, users may be encouraged to feed raw logs into automation pipelines, increasing the risk of inadvertent sensitive-data exposure, retention, or over-broad access inside the agent environment.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The "When to Use" section includes vague trigger phrases like "what's failing", "improve my agent", and "check system health" without boundaries or confirmation requirements. In agentic systems, such broad matching can cause the skill to activate during routine troubleshooting or planning dialogue, unnecessarily exposing internal logs or causing unreviewed self-improvement recommendations to be generated from sensitive operational data.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The file’s header describes the skill in English, but the action descriptions and section labels switch to Chinese for key operational details. Because the skill does not offer a language/locale option or explain that it is intended for Chinese-speaking users, this creates a natural-language locale policy issue.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.