Back to skill

Security audit

QClaw Self-Evolver

Security checks for vulnerabilities and agentic risk

Overview

This skill is openly a self-evolution system, but it installs persistent scheduled agent activity and can turn stored user corrections into future skill or agent-instruction changes without strong approval boundaries.

Install only if you intentionally want an agent that keeps local learning records, creates candidate skills, and runs a scheduled self-evolution scan. Review generated skills before enabling them, avoid storing secrets in corrections or error logs, and disable or remove the SEA-Evolution-Scan cron job unless you explicitly want recurring background agent activity.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • System PersistenceInstalls backdoors, hooks, services, or scheduled tasks that survive the run
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T06 · System Persistence

Error
Location
scripts/install_self_evolver.py:97
Finding

Recurring Agent Task Creates Cross-Session Persistence

Content
View full analysis

Vulnerability Details

File Location: scripts/install_self_evolver.py:97-131
Vulnerability Type: Scheduled cross-session agent execution
Risk Level: High

Vulnerable Code

python
def register_cron():
    """注册每3天运行 SEA 进化的定时任务"""
    try:
        from openclaw.cron import CronManager
        import os

        # 检测是否是 openclaw 环境
        try:
            import openclaw
        except ImportError:
            print("  ℹ️  openclaw 模块不可用,跳过 cron 注册")
            print("     请手动添加 cron 任务:")
            print("     openclaw cron add --name 'SEA-Evolution' --every 3days")
            return

        cm = CronManager()
        job = {
            "name": "SEA-Evolution-Scan",
            "schedule": {"kind": "every", "everyMs": 3 * 24 * 60 * 60 * 1000},
            "payload": {
                "kind": "agentTurn",
                "message": "运行自我进化:python ~/.qclaw/workspace/scripts/sea_evolve.py --quiet",
            },
            "sessionTarget": "isolated",
            "enabled": True,
        }
        # 检查是否已存在
        existing = [j for j in cm.list() if j.get("name") == "SEA-Evolution-Scan"]
        if existing:
            print("  ℹ️  cron 任务已存在,跳过")
        else:
            cm.add(job)
            print("  ✅ cron 任务已注册(每3天运行)")
    except Exception as e:
        print(f"  ℹ️  cron 注册失败(非致命):{e}")
        print("     请手动添加:每3天运行 python ~/.qclaw/workspace/scripts/sea_evolve.py --quiet")

Technical Analysis

Installation creates an enabled OpenClaw scheduled task that initiates an isolated agent turn every three days. The task survives the installation session and asks the agent to execute a Python script from the writable user workspace.

Although the recurring behavior is disclosed in the documentation, the installer registers it automatically without an explicit confirmation immediately before persistence is established. The scheduled payloa ...[truncated 1157 chars]

Remediation
View remediation

Remediation Suggestions

  • Do not register the recurring task by default. Require explicit, informed confirmation immediately before calling cm.add().
  • Offer manual execution as the default installation mode.
  • Execute an immutable or integrity-verified script instead of trusting a mutable workspace path.
  • Verify ownership and restrictive permissions on the workspace and scheduled script.
  • Store and verify a cryptographic digest before each scheduled execution.
  • Provide a documented uninstall operation that removes SEA-Evolution-Scan.
  • Display the exact schedule, payload, target session, and removal command before registration.
  • Consider requiring approval for each evolution cycle rather than launching an autonomous agent turn.

T01 · Skill Instruction Hijacking

Error
Location
scripts/skill_evolution.py:52
Finding

Untrusted Input Is Written Verbatim into Persistent Skill Instructions

Content
View full analysis

Vulnerability Details

File Location: scripts/skill_evolution.py:52-108
Vulnerability Type: Persistent instruction and Markdown injection
Risk Level: High

Vulnerable Code

python
def generate_skill(name: str, description: str, trigger_phrases: list,
                   skill_type: str = "task-automation") -> str:
    """
    生成一个新的 skill SKILL.md 文件
    返回生成的路径
    """
    skill_dir = os.path.join(SKILLS_DIR, f"{_slugify(name)}-skill")
    os.makedirs(skill_dir, exist_ok=True)

    skill_md = f"""# {name}

**类型:** {skill_type}
**生成时间:** {datetime.now().strftime("%Y-%m-%d %H:%M")}
**触发词:** {', '.join(trigger_phrases) if trigger_phrases else '(待添加)'}
**状态:** 实验性(待验证效果)

---

## 描述

{description}

---

## 触发条件

> 当用户提到以下内容时,触发本技能:

{"".join(f"- {t}\n" for t in trigger_phrases)}

---

## 使用方法

(待填写:具体使用步骤)

---

## 注意事项

- 本技能为自动生成,效果待验证
- 使用 3 次后评估效果
- 如需改进请联系维护者

---

*本技能由 self-evolver 自动生成 · {datetime.now().strftime("%Y-%m-%d")}*
"""
    skill_path = os.path.join(skill_dir, "SKILL.md")
    with open(skill_path, "w", encoding="utf-8") as f:
        f.write(skill_md)

    # 更新 references/index.md(方便查找)
    index_path = os.path.join(SKILLS_DIR, "_auto_generated_index.md")
    with open(index_path, "a", encoding="utf-8") as f:
        f.write(f"- [{name}]({_slugify(name)}-skill/SKILL.md) — {description[:50]}\n")

    return skill_path

Technical Analysis

The name, description, skill_type, and trigger_phrases values originate from command-line arguments and are interpolated directly into a generated SKILL.md file. No validation, structural escaping, instruction filtering, or approval boundary is applied.

An attacker able to influence these values can include line breaks, Markdown headings, quoted directives, or instruction-like content. These values are therefore not confined to descriptive metadata: they can alter the structure a ...[truncated 1626 chars]

Remediation
View remediation

Remediation Suggestions

  • Treat every generated field as untrusted data.
  • Reject or normalize multiline input, headings, block quotes, code fences, links, and other Markdown control syntax.
  • Define strict length, character, and format allowlists for names, types, descriptions, and triggers.
  • Serialize metadata in a strict schema and render it through an escaping template rather than direct string interpolation.
  • Detect and reject instruction-control phrases that attempt to override higher-level rules.
  • Store generated candidates outside directories that the agent automatically loads.
  • Require human review and explicit approval before moving a candidate into the active skills/ directory.
  • Clearly delimit untrusted descriptive content and instruct the Skill loader never to execute instructions found inside data fields.
  • Apply the same escaping rules when writing _auto_generated_index.md.
  • Maintain provenance and an audit log identifying which user input produced each generated Skill.

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:80
Finding

Skill Directs Persistent Agent Instructions to Be Updated from Learning Records

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:80-86
Vulnerability Type: Persistent agent-memory poisoning workflow
Risk Level: High

Vulnerable Instruction

markdown
**Evolve(进化):**
- 对连续错误 → 更新系统提示词 / AGENTS.md
- 对重复任务 → 触发技能自进化生成新 skill
- 生成进化报告

Technical Analysis

The Skill specification directs the agent to update the system prompt or AGENTS.md in response to repeated errors. Correction records can contain user-originated content, but the documented process does not define a trust boundary, validation policy, approval requirement, constrained rule format, or rollback mechanism before converting that content into persistent agent instructions.

The bundled Python implementation does not directly perform the stated AGENTS.md or system-prompt modification. However, SKILL.md is itself the behavioral instruction interface presented to an agent. An agent implementing the documented evolution process could therefore promote poisoned correction content into persistent governing rules.

Attack Path

  1. An attacker submits crafted correction text during interactions with the agent.
  2. The correction is placed into the learning records, potentially multiple times so it appears to be a repeated error pattern.
  3. The evolution workflow identifies the repeated content as an improvement candidate.
  4. Following SKILL.md, the agent updates its system prompt or AGENTS.md.
  5. The malicious rule then affects future sessions that consume the modified persistent instructions.

Impact Assessment

The potential scope includes persistent alteration of future agent decisions, safety constraints, task priorities, and tool-use behavior. Any concrete system impact remains bounded by the agent's existing permissions. The repository does not contain code that directly edits AGENTS.md, so exploitation depends on an agent following and implementing the documented instruction.

Remediation
View remediation

Remediation Suggestions

  • Remove instructions that permit automatic modification of system prompts or AGENTS.md.
  • Generate a proposed patch or recommendation report instead of applying changes directly.
  • Require explicit workspace-owner approval for every persistent instruction change.
  • Permit only narrowly defined rule formats and reject arbitrary free-form instructions from learning records.
  • Preserve source provenance so reviewers can trace every proposed rule to its originating interaction.
  • Add deduplication and trust scoring; repetition alone must not establish that a correction is trustworthy.
  • Maintain versioned backups, change logs, and a reliable rollback procedure.
  • Separate untrusted learning data from trusted policy and instruction files.
  • Prevent lower-trust users or sessions from proposing changes to system-level behavior.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (25)

Self-Modification

High
Category
Rogue Agent
Confidence
98% confidence
Finding

The skill explicitly installs a self-evolution capability that creates persistent state, schedules recurring scans, and enables future autonomous behavior changes. In context, this is more dangerous than a normal automation skill because it authorizes self-modification and persistence mechanisms that can alter agent behavior over time without bounded review.

Content

Scanner excerpt · SKILL.md (reported line 23)May include surrounding context.

md
## 触发条件

> 当用户说「帮我安装自我进化」或「安装 self-evolver」或「你要有自我进化能力」时,触发本 skill。

安装后 AI 会:
1. 创建 `.learnings/` 目录结构

Self-Modification

High
Category
Rogue Agent
Confidence
98% confidence
Finding

The installation command for self-evolver operationalizes the self-modification design by adding a skill whose purpose is autonomous learning and skill generation. This is high risk because it creates a persistent mechanism for behavior expansion and file generation from historical interactions, which can amplify errors or adversarial inputs across future runs.

Content

Scanner excerpt · SKILL.md (reported line 37)May include surrounding context.

1. 安装(首次)

text
skillhub_install install_skill self-evolver

或在 OpenClaw 中对 AI 说:

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 46)May include surrounding context.

md
- 生成 `SKILL.md` 覆盖本文件

Self-Modification

High
Category
Rogue Agent
Confidence
90% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · scripts/install_self_evolver.py (reported line 12)May include surrounding context.

python
4. 注册 cron 定时任务(每3天运行 SEA 进化扫描)

使用方式:
    python install_self_evolver.py
"""

import os

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The script automatically registers a recurring cron job that will trigger future agent execution every three days, creating persistence without explicit user approval. Scheduled execution is more dangerous than one-time file creation because it can repeatedly run evolving code later, potentially after the user forgets it was installed or after the copied scripts change.

Content

No source excerpt is available for this finding.

Self-Modification

High
Category
Rogue Agent
Confidence
90% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · scripts/skill_evolution.py (reported line 99)May include surrounding context.

python
---

*本技能由 self-evolver 自动生成 · {datetime.now().strftime("%Y-%m-%d")}*
"""
    skill_path = os.path.join(skill_dir, "SKILL.md")
    with open(skill_path, "w", encoding="utf-8") as f:

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill is designed to persist user corrections and task-history-derived skills into long-term local storage. Persistent retention of interaction history is dangerous because it can capture sensitive content, create privacy liabilities, and later influence future behavior in ways the user did not intend.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The activation instructions are written as fixed Chinese phrases, and later usage examples and trigger keywords are also predominantly Chinese, without stating that the skill is region-specific or offering users a language alternative. This creates a natural-language locale constraint that is not justified or opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill registers a recurring cron job every 3 days but does not clearly warn the user that it is installing persistent scheduled execution. Background scheduled tasks expand the attack surface by enabling ongoing processing and future actions without an interactive prompt.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill registers a recurring cron job every 3 days but does not clearly warn the user that it is installing persistent scheduled execution. Background scheduled tasks expand the attack surface by enabling ongoing processing and future actions without an interactive prompt.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill automatically analyzes conversation/task history and logs corrections, but the description lacks a privacy warning or consent mechanism. This is risky because sensitive user inputs, mistakes, and requests may be retained and mined persistently without the user's informed approval.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill claims to implement local learning/evolution behavior, but its documented Evolve step explicitly includes updating AGENTS.md and system prompts, which changes higher-trust agent behavior outside simple note-taking. That is dangerous because user corrections or noisy history could be transformed into persistent instruction changes, effectively enabling prompt/configuration drift and unauthorized self-modification.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The workflow instructs the agent to immediately write detected user corrections to pending storage and later merge them into durable records. Immediate persistence of interpreted user intent is risky because misclassification, sensitive disclosures, or adversarial prompt content can be stored and later treated as trusted learning material.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The documented directory structure stores corrections, errors, feature requests, and pending user inputs in persistent local files. This expands the local sensitive-data footprint and can expose private user content to other local processes, future skills, or anyone with filesystem access.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The script persistently stores user-supplied corrections, command failures, and related context in plaintext markdown files under a workspace directory. If users include secrets, credentials, internal prompts, or sensitive operational context in these messages, the tool creates a durable local data-retention surface that can later be exposed through backup systems, sync tools, shared home directories, or other local access.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The module docstring states that detected corrections are immediately written to pending.md and later merged into LEARNINGS.md, which changes files under the workspace. Although the script prints status after some CLI actions, there is no up-front warning or confirmation around the automatic write/merge behavior described here, including the clearing of pending.md during flush.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The correction-detection design is based on broad natural-language keywords and semantically encourages capture of arbitrary user messages that contain phrases like 'wrong' or 'stop'. In a larger agent workflow, this can cause overcollection of unrelated or sensitive user text and silently convert conversational content into persistent logs, increasing privacy and data leakage risk beyond the minimum needed for the feature.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

Natural-language strings throughout the file, including the top-level description and usage guidance, are fixed in Chinese. Under the policy, forcing a specific language without user opt-in is a locale/language policy violation unless the constraint is clearly documented as region-specific or optional.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The installer performs persistent filesystem changes in the user's workspace by copying executable scripts and creating new directories/files without any prior consent, dry-run, or interactive confirmation. In an agent-skill context, silent installation behavior increases the risk of unexpected persistence and makes it easier to plant or replace automation components the user did not explicitly approve.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This code file contains its primary description, usage instructions, and user-facing output entirely in Chinese, with no indication that language is selectable or optional. Under the stated policy, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The natural-language instructions, trigger descriptions, CLI help text, and generated SKILL.md template all assume Chinese output and interaction. Under the language/locale policy, forcing a single language without opt-in or justification is a policy concern.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This code creates and writes a new SKILL.md file, which is a user-data-affecting file write operation. Although the script prints a success message afterward, there is no confirmation prompt before writing and the docstring/commentary does not warn that running with --name will create or overwrite skill content on disk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

In addition to generating the requested skill file, the script silently appends an entry to _auto_generated_index.md. This is a second file write with persistent side effects, but the code gives no advance notice that invoking skill generation will also mutate the index.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The module docstring and user-facing CLI output are written only in Chinese, which indicates a fixed language choice in the skill's natural-language interface. The file does not offer a language/locale option or explain a justified region-specific constraint, so this may violate language-choice policy.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

All user-facing instructions and prompts in the file are in Chinese, and the script does not provide any opt-in or alternative language handling. This can violate language/locale policy when a skill forces a specific language without giving the user a choice or documenting a justified regional constraint.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.