Back to skill

Security audit

Continuous Evolution

Security checks for vulnerabilities and agentic risk

Overview

The skill is not overtly malicious, but it automatically stores task details and creates follow-up queued tasks in shared OpenClaw workspace locations without enough user control or data minimization.

Install only if you are comfortable with task descriptions, execution results, and audit scores being written into shared OpenClaw memory files and with low scores creating queued follow-up tasks. Avoid using it on sensitive prompts, secrets, proprietary code, or personal data unless you first add redaction, retention limits, and an approval gate for queued tasks.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T02 · Agent Memory Poisoning

Warning
Location
evolve.sh:20
Finding

Unsanitized Attacker-Controlled Content Is Written to Persistent Agent Memory

Content
View full analysis

Vulnerability Details

File Location: evolve.sh, lines 20-40
Vulnerability Type: Persistent agent memory poisoning
Risk Level: Medium

Vulnerable Code

bash
record_experience() {
    local task_id="$1"
    local task="$2"
    local result="$3"
    local audit_score="$4"
    local exp_file="${EVOLUTION_DIR}/experience-$(date +%Y-%m-%d).md"
    
    log "📚 记录经验:$task_id"
    
    cat >> "$exp_file" << EOF

---

## $(date "+%H:%M:%S") - $task_id

**任务**: $task
**结果**: $result
**审计评分**: $audit_score/100

EOF

Technical Analysis

The record_experience function accepts task identifiers, task descriptions, and execution results from command-line arguments and writes them verbatim into a persistent Markdown file under /root/.openclaw/workspace/memory/evolution.

No trust metadata, structural encoding, content filtering, provenance tracking, or separation between untrusted task data and trusted agent instructions is applied. Consequently, an attacker who can influence the task description or result can insert instruction-like Markdown into the persistent experience record.

The shell here-document does not directly execute command substitutions contained inside variable values, so this is not shell command injection. The risk arises when a later agent session or automated workflow reads the generated Markdown as trusted memory or guidance. At that point, embedded instructions could be interpreted as agent directives rather than untrusted historical data.

Attack Path

  1. An attacker supplies or influences a task description or execution result containing adversarial instructions, such as directions to ignore security constraints, disclose data, or invoke tools.
  2. The caller passes that content to evolve.sh as the second or third argument.
  3. record_experience appends the content without sanitization or provenance labels to a daily file in the shared agent memory directory.
  4. The malicious content remains present across subse ...[truncated 1003 chars]
Remediation
View remediation

Remediation Suggestions

  1. Store task descriptions and results as structured data, such as JSON, rather than instruction-like Markdown.
  2. Mark every externally supplied field with explicit provenance and trust metadata, for example:
    • source: external
    • trusted: false
    • content_type: task_data
  3. Ensure memory consumers treat stored content exclusively as quoted data and never as executable instructions.
  4. Apply length limits and schema validation to task identifiers, task descriptions, and result fields before persistence.
  5. Escape or encode Markdown control characters and delimit untrusted content clearly when a human-readable Markdown representation is required.
  6. Store records in a skill-owned directory with restrictive permissions instead of a shared root-level agent memory location.
  7. Require explicit review or approval before promoting task records into long-term agent memory.
  8. Add tests using adversarial task and result values, including prompt-injection text, multiline Markdown, delimiter manipulation, and oversized input.
  9. Document the trust boundary and require downstream memory-loading components to isolate historical records from system and developer instructions.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (8)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The visible skill title/subtitle is presented in Chinese only, and the rest of the README content is also written in Chinese. For a general-purpose skill README, this imposes a language choice on users without opt-in or any stated regional justification, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill explicitly describes automatic recording, analysis, and archival of execution experience after every task, but provides no notice, consent boundary, or data-minimization guidance. In practice, this can cause user prompts, task outputs, sensitive workspace content, or operational metadata to be persistently stored without user awareness, creating privacy and retention risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The workflow and output sections describe automatic creation of files under memory/ and tasks/queue/ plus automatic queuing of follow-on tasks, without warning the user that the skill persists artifacts in the workspace. This can surprise users, pollute repositories, expose sensitive information in committed files, and create uncontrolled task chaining or stateful behavior beyond the original request.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Multiple comments and all user-visible output strings, including the usage message and operational logs, are written in Chinese only. This imposes a fixed language on users without opt-in or documented locale justification, which matches the language policy violation criteria.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The script persists raw task descriptions, execution results, and audit scores to a dated markdown file under a memory directory with no sanitization, minimization, retention control, or user consent mechanism. In an agent skill context, tasks and results can contain sensitive prompts, secrets, proprietary code, or personal data, so automatic long-term storage increases confidentiality and privacy risk if those files are later read, indexed, or exfiltrated.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

This markdown file uses Chinese throughout (for example, the title and section contents) and does not indicate that the skill is region-specific or that users may choose another language. Under the policy rule for natural-language violations, forcing a specific language without user opt-in can be a violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The file's natural-language instructions and labels are entirely in Chinese, with no indication that users may choose another language. Under the policy rule, forcing a specific language without opt-in can be a locale-policy issue unless the restriction is explicitly justified.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The script automatically creates new queued task files based on low audit scores, causing autonomous persistence and follow-up work generation without an approval gate. In this skill context that is less about classic code execution and more about unsafe agent autonomy: it can create unexpected tasks, amplify bad inputs, and consume resources or trigger unintended downstream processing.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.