Back to skill

Security audit

Self Evolution-v2

Security checks for vulnerabilities and agentic risk

Overview

This skill is openly aimed at persistent agent self-review, but it changes long-term agent instructions and memory behavior with broad automatic triggers and limited user control.

Install only if you intentionally want this skill to change your OpenClaw workspace's persistent behavior. Review the exact SOUL.md block first, keep a backup, and be prepared to remove the inserted rules and .learnings files if you do not want future sessions to read prior memory or store inferred feedback automatically.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (2)

T01 · Skill Instruction Hijacking

Error
Location
scripts/inject_to_soul.sh:52
Finding

Persistent Agent Instruction Hijacking Through SOUL.md Modification

Content
View full analysis
> "$SOUL_FILE" << 'EOF' --- ## 🚨 自我进化规则(Self-Evolution Skill - 自动加载) > 此部分由 self-evolution 技能自动注入,每次会话自动携带,无需调用。 ### 🚀 新会话自动回顾(/new 时自动执行) **触发时机:** 当输入 `/new` 开始新会话时 **AI会自动执行:** 1. 读取 memory/YYYY-MM-DD.md(昨天的记录) 2. 扫描上一次对话中的:老板纠正了什么、我说了哪些不确定的话、老板哪些事没回音、哪些任务没交付完 3. 把发现的问题写入 .learnings/LEARNINGS.md 4. 把重要教训同步到 SOUL.md 5. 继续当前任务,不丢失上下文 ### P0 最高优先:不确定 → 立即查证 **触发词警报:** 发现自己想说"可能"、"大概"、"应该是"、"不确定"的时候 → 立刻停 → 动手查 ``` The injected block continues through line 163 and defines response policies, proactive behavior, memory-writing triggers, communication rules, self-assessment behavior, and persistent core principles. ### Technical Analysis The script appends Skill-controlled instructions directly to: ```bash SOUL_FILE="$HOME/.openclaw/workspace/SOUL.md" ``` `SOUL.md` is described by the Skill as an automatically loaded Agent instruction file. The injected content explicitly declares that it will be carried into every session without requiring the Skill to be invoked again. It also declares part of its behavior as the highest priority. This violates instruction-boundary isolation. A Skill should provide behavior only while it is explicitly invoked and should not silently promote its own instructions into a globally loaded identity or policy file. By writing to `SOUL.md`, the Skill converts locally scoped Skill content into persistent Agent-level instructions that can influence unrelated future requests. The injected rules affect more than the stated installation task. They alter when the Agent may respond, what it must record, when it should proactively contact the user, how it interprets user sentiment, and how it handles future sessions. ### Attack ...[truncated 1293 chars]
Remediation
View remediation

T02 · Agent Memory Poisoning

Error
Location
scripts/inject_to_soul.sh:64
Finding

Persistent Memory Poisoning From Unverified Conversation Inferences

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (25)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的是高层次的 agent 元能力:自我审视、持续进化、类似 Hermes Agent 的核心能力。但代码仅实现了一个非常具体且有限的功能——扫描一句话中的不确定措辞,按预设关键词规则分类为 L1/L2/L3,并打印建议。这属于简单的规则式置信度表达检查工具,不涉及自我反思、长期记忆、学习、能力演化、自动改进或任何复杂 agent 行为。因此其主要目的与声明严重不符,应判定为描述与行为不匹配。

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
79% confidence
Finding

The skill instructions, trigger phrases, example outputs, and policy language are entirely specified in Chinese and assume Chinese conversational markers such as "老板说" and Chinese trigger words. There is no indication that users may opt into another language or that the skill is intentionally restricted to a Chinese-language environment.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill says it will automatically write to SOUL.md and create template files after a simple activation phrase, without an explicit confirmation step, scope preview, or change summary. Any skill that modifies persistent agent configuration or workspace files implicitly can cause unintended configuration drift, overwrite local customizations, or be abused to inject durable behavioral changes.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The activation instruction is a short natural-language phrase with no scope constraints, exclusions, or confirmation semantics. Broad activation phrases increase the risk of accidental triggering, prompt-injection-style reapplication, or silent enabling of persistent behaviors in contexts where the user did not intend to modify memory, policies, or files.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill instructs the agent to automatically read prior memory files at new-session start and surface unresolved lessons and user-related context. This creates a cross-session data propagation channel that may expose prior sensitive details, carry forward stale or private context without consent, and undermine user expectations that a new session starts cleanly.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill repeatedly requires immediate logging of user feedback, corrections, and task outcomes into persistent files. Persistent recording of interaction details can accumulate sensitive operational context, user preferences, mistakes, or project data without clear minimization, retention limits, or consent, increasing privacy and data exposure risk.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill defines many broad triggers based on ordinary conversational phrases like '好了', '继续', '算了', and inferred user behavior. This can cause unintended activation of logging, state changes, or behavioral adaptations from ambiguous or routine chat, making the agent persist or act on inferred signals the user never intended as control inputs.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill mandates detecting and recording implicit behavioral signals such as repeated questions, terse replies, inferred dissatisfaction, or emotional cues, then persisting them as lessons or errors. This is more dangerous than ordinary task logging because it formalizes surveillance-like profiling of user behavior and sentiment, which can misclassify users, retain sensitive inferences, and influence future agent behavior in opaque ways.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The document is entirely written in Chinese and provides instructions only in that language, with no indication that language selection is optional or that the template is intended solely for a Chinese-language environment. This creates a natural-language locale policy concern because it implicitly mandates a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

All user-facing instructions in this template are written in Chinese, and the trigger phrase examples are also language-specific, with no indication that other languages are supported or that Chinese is required for a documented regional purpose. This can violate language/locale policy when a skill imposes a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This markdown file defines auto-trigger rules such as '遇到全新领域任务' and '老板问了一个我不知道的概念' without clear boundaries, exclusions, or negative examples. These conditions are broad enough to match many normal interactions, which could cause unintended or excessive invocation of the skill/template.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The template content is entirely in Chinese and implicitly requires the agent to operate in that language/locale, without any user preference check or fallback. This can cause the agent to ignore user language expectations, reduce usability, and potentially lead to misunderstandings in task tracking or status communication.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The document title and all template content are written exclusively in Chinese, which creates a natural-language policy concern if the organization requires language choice or user opt-in before forcing a specific locale. There is no indication anywhere in the file that another language is allowed or that Chinese is optional.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The file is entirely written in Chinese and provides the invocation phrase only in Chinese, without indicating that other languages are supported or that the language choice is optional. This can constitute a language/locale policy issue because the skill appears to require a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documentation explicitly tells the AI to automatically write evolution rules into SOUL.md and create learning template files in the user's workspace, but it does not warn the user that persistent configuration files will be modified. Because these are long-lived agent state files, silent or implicit modification can change future agent behavior and create hard-to-audit persistence, making the skill more dangerous in the context of a self-modifying or self-evolving agent.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This code’s user-facing description and usage are entirely in Chinese, and all runtime prompts and classifications are also hard-coded in Chinese. Because the file provides no opt-in, language selection, or justification that the tool is region-specific, it imposes a specific language/locale policy on users.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The script's user-facing title, usage notes, and operational messages are entirely in Chinese, and the injected SOUL.md content is also written only in Chinese. There is no indication that the skill is region-specific or that users can opt into another language, which can violate language/locale choice policy.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script proceeds to append behavioral rules into SOUL.md without an explicit warning or interactive confirmation at the point of modification. Silent persistent modification of an agent's instruction base undermines informed consent and makes it easier for users to install behavior-changing logic without appreciating long-term effects.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The script persistently modifies the agent's core SOUL.md by appending auto-loaded behavioral rules that affect all future sessions. This crosses from optional self-review into durable policy injection, creating a supply-chain style trust risk where installing a skill silently changes the agent's governing behavior beyond the current invocation.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill description emphasizes self-inspection and evolution, but the implementation actually rewrites the session-governing SOUL.md with automatically loaded rules. That mismatch is dangerous because users may consent to a reflective aid without understanding they are installing persistent instruction-layer modifications that can influence future decisions and data handling.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The injected rules instruct the AI to automatically read prior conversation artifacts and persist derived content into .learnings and SOUL.md across sessions. This creates uncontrolled retention and propagation of user-provided data, including potentially sensitive instructions, preferences, mistakes, or confidential context, without any filtering, minimization, or consent checks.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The auto-memory section tells the agent to record user corrections, preferences, dissatisfaction signals, and task outcomes into persistent files automatically. Persisting this behavioral and preference data without sensitivity screening or consent can capture personal, confidential, or manipulative content and make it available to future sessions in ways the user did not intend.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The script's comments and all user-visible prompts are in Chinese, which imposes a specific language on users without any opt-in or indication that the skill is intentionally limited to a Chinese-speaking context. This matches the policy category for language/locale constraints that are not optional or justified.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · scripts/self_check.sh (reported line 56)May include surrounding context.

sh
fi

echo ""
echo "🎯 如需完整分析,运行: cat ~/.openclaw/skills/self-evolution/SKILL.md"

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

This markdown file contains user-facing natural-language instructions exclusively in Chinese, which can constitute a language policy violation when no opt-in or alternative language is offered. The file does not indicate that the template is intended only for a Chinese-speaking or region-specific context.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.