Back to skill

Security audit

Self-Improve

Security checks for vulnerabilities and agentic risk

Overview

This self-improvement skill is mostly disclosed, but it needs Review because it can repeatedly read and share agent memories/logs and includes unsafe file write/delete utilities.

Review before installing. Do not approve the cron entry unless you are comfortable with recurring cross-agent memory/log aggregation into shared storage. Restrict workspace_root to explicitly approved agents, add redaction and retention controls, and fix the path traversal issues in the CLI scripts before exposing them to untrusted input or automation.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (6)

T05 · Unauthorized Access and Privilege Escalation

Error
Location
scripts/setup.mjs:380
Finding

Cross-Agent Memory Collection Violates Agent Isolation Boundaries

Content
View full analysis
只有固化到系统文件时需要确认,其他全自动。 --- ## [P-001] 添加 Cron 定时任务 - **来源:** self-improve 安装 - **建议修改:** 在 openclaw.json 的 cron 中添加: \`\`\`json { "name": "self-improve", "schedule": { "kind": "cron", "expr": "${CRON_EXPR}", "tz": "${OWNER_TZ}" }, "sessionTarget": "isolated", "payload": { "kind": "agentTurn", "message": "你是 Self-Improve 系统的执行者。请按以下步骤执行:\\n1. 读取 ${ROOT.replace(/\\/g, '\\\\\\\\')}\\\\ENGINE.md 了解完整流程\\n2. 读取 ${ROOT.replace(/\\/g, '\\\\\\\\')}\\\\config.yaml 了解模块配置\\n3. ${workspaceScanInstruction}\\n4. 按执行顺序运行所有已启用模块\\n5. ${notificationInstruction}\\n6. 记录运行日志到 run-log.jsonl", "model": "${CRON_MODEL}" } } \`\`\` `; ``` The associated documentation explicitly establishes the shared-access behavior: ```markdown **All agents share this system.** When self-improve runs: 1. Scan `/path/to/self-improve/` (system-level, shared by all agents) 2. Read each agent's session logs (if accessible) 3. Write results to shared location, readable by all agents ``` ### Technical Analysis The setup script generates a recurring Agent task whose instructions direct the executing Agent to scan every Agent memory directory under the configured workspace. The results are intended to be written into a shared location readable by all Agents. Filesystem accessib ...[truncated 1657 chars]
Remediation
View remediation

T02 · Agent Memory Poisoning

Error
Location
scripts/improve.mjs:127
Finding

Untrusted Feedback Can Be Converted into Persistent Agent Instructions

Content
View full analysis
r.hint).map(r => r.hint); const freq = {}; hints.forEach(h => { freq[h] = (freq[h] || 0) + 1; }); const topIssues = Object.entries(freq).sort((a, b) => b[1] - a[1]).filter(([_, c]) => c >= 2); let promptAdditions = ''; if (topIssues.length > 0) { promptAdditions = topIssues.map(([hint, count], i) => ` ${i + 1}. 注意: ${hint} (反馈 ${count} 次)`).join('\n'); } // 提取成功案例 const goodExamples = records.filter(r => r.score > 0 && r.hint).slice(-3); let examplesYaml = ''; if (goodExamples.length > 0) { examplesYaml = goodExamples.map(r => ` - "${r.hint || 'good'}"`).join('\n'); } ``` The resulting text is written into persistent prompt material: ```javascript if (promptAdditions) { if (content.includes('prompt_additions:')) { content = content.replace(/prompt_additions: \|[\s\S]*?(?=\n\w)/, `prompt_additions: |\n${promptAdditions}\n`); } else { content += `\nprompt_additions: |\n${promptAdditions}\n`; } } writeFileSync(filePath, content, 'utf-8'); ``` Feedback is initially persisted verbatim by `feedback.mjs`: ```javascript const entry = `\n## ${today()} ${new Date().toTimeString().slice(0, 5)} - **agent:** ${record.agent} - **信号:** ${record.type === 'explicit' ? '用户纠正' : record.type} - **内容:** "${record.hint}" - **次数:** 1/3 - **状态:** ⏳ 待观察 `; ``` Arbitrary content can also be inserted directly into HOT memory: ```javascript const level = opts.level || 'hot'; const content = opts.content; const ts = new Date().toISOString().split('T')[0]; const entry = `- ${content} ($ ...[truncated 2314 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/module.mjs:112
Finding

Unchecked Module Name Allows Recursive Directory Deletion Outside the Module Directory

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/memory-ops.mjs:167
Finding

Unchecked Domain and Project Names Permit Writes Outside Memory Directories

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/improve.mjs:45
Finding

Unchecked Task Name Permits Skill File Reads and Writes Outside the Skills Directory

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/run-all.mjs:60
Finding

Module Runner Records Successful Execution Without Invoking Modules

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (51)

Ssd 3

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

These instructions create a natural-language exfiltration path: read each agent's session logs if accessible, then write distilled results to a shared location readable by all agents. That enables cross-agent leakage of secrets, private user conversations, internal reasoning artifacts, and other sensitive data without isolation boundaries.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The document states that system file modifications require approval, but later provides direct instructions to add content to AGENTS.md, which is listed as requiring confirmation. This contradiction weakens the safety boundary and can normalize unauthorized modification of trusted instruction files.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The onboarding section instructs adding content to new agents' AGENTS.md files, which is effectively a mechanism for propagating behavioral instructions into other agents. That exceeds passive self-improvement and becomes cross-agent instruction injection, potentially altering agent behavior without a clear authorization boundary.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The remove functionality provides destructive capability that deletes module directories from disk, which is a powerful operation not clearly required for the stated self-improvement purpose. In the context of an agent skill that may run with broad filesystem access, this can be exploited or accidentally triggered to remove components and impair or subvert the framework.

Content

No source excerpt is available for this finding.

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · scripts/feedback.mjs (reported line 1)May include surrounding context.

js
#!/usr/bin/env node
/**
 * 系统状态查看工具
 *

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · scripts/improve.mjs (reported line 1)May include surrounding context.

js
#!/usr/bin/env node
/**
 * 系统状态查看工具
 *

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · scripts/memory-ops.mjs (reported line 1)May include surrounding context.

js
#!/usr/bin/env node
/**
 * 系统状态查看工具
 *

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · scripts/module.mjs (reported line 1)May include surrounding context.

js
#!/usr/bin/env node
/**
 * 系统状态查看工具
 *

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · scripts/report.mjs (reported line 1)May include surrounding context.

js
#!/usr/bin/env node
/**
 * 系统状态查看工具
 *

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · scripts/run-all.mjs (reported line 1)May include surrounding context.

js
#!/usr/bin/env node
/**
 * 系统状态查看工具
 *

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · scripts/status.mjs (reported line 1)May include surrounding context.

js
#!/usr/bin/env node
/**
 * 系统状态查看工具
 *

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The approval model claims only system-file solidification needs confirmation, but later sections permit automatic creation of files and directories. This inconsistency can mislead users about what the skill may change automatically and weakens informed consent around persistent modifications.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The document authorizes automatic writes to broad data locations such as self-improve data and learned knowledge stores, rather than limiting behavior to approval-gated proposals. This increases the attack surface for persistent prompt/data poisoning, accidental propagation of bad rules, and unauthorized state changes without meaningful user review.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
83% confidence
Finding

The skill grants the agent autonomous authority to create directories and files without prior approval. While likely intended for convenience rather than abuse, in this context it materially increases the risk of unauthorized persistent state changes, prompt injection persistence, and silent expansion of writable surface area.

Content

Scanner excerpt · ENGINE.md (reported line 62)May include surrounding context.

md
## Layer Creation Rules

**Agents can automatically create new directories/files without asking in advance.**

### themes/ Directory (Extensible)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The engine states that processing is fully automatic and can create new files/directories without a clear warning to the user about ongoing data-modifying behavior. This can cause unnoticed persistent changes, knowledge pollution, or storage of sensitive derived content in places the user does not expect.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The engine explicitly instructs scanning a shared system area, reading other agents' session logs, and publishing results to a shared location readable by all agents. That goes well beyond self-improvement proposal generation and creates a direct cross-agent data access and redistribution channel, which can expose sensitive prompts, user data, or secrets from unrelated sessions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill describes automatic access to other agents' logs and shared publication of results without any prominent privacy warning, consent flow, or data handling constraints. In context, this is especially dangerous because session logs commonly contain private user content, credentials, internal prompts, and sensitive operational history.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The runtime explicitly directs the agent to write and update multiple local files at each step, but it does not require any user-facing notice, consent, or confirmation before modifying state. In an autonomous skill that runs on a schedule and survives crashes via checkpoints, silent persistent writes increase the risk of unexpected state changes, clutter, or overwriting project artifacts without the operator realizing it.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This section describes routing content to multiple destinations, including paths outside the self-improve directory and eventual publication to a blog platform, without requiring explicit user approval or destination validation. That creates a real risk of unintended data propagation, disclosure of sensitive internal context, or uncontrolled writes to external/shared locations, especially because the system reuses prior logs and distilled material automatically.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manual trigger phrase "learn and improve" is broad enough to overlap with ordinary conversational requests, which can cause the skill to activate unintentionally. In a self-modifying or system-improving framework, accidental invocation is more dangerous because it may start scanning memory, generating proposals, or influencing system files without the user explicitly intending to launch this higher-risk workflow.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The document frames the system as execution-quality improvement, but it explicitly authorizes broad downstream outputs such as team-wide knowledge sharing and blog drafts. That expands the data flow beyond narrowly scoped self-improvement and creates a real risk of repurposing internal agent data into wider artifacts without clear consent and minimization controls.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The daily-work logging model records corrections, praise, important tasks, and reusable rules into persistent files, but the surrounding documentation does not present a prominent user-facing warning about the scope and privacy consequences of this collection. This makes silent accumulation of work artifacts more likely and undermines informed consent.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · ENGINE.md (reported line 267)May include surrounding context.

md
- Applied 3 times/7 days → Elevate to HOT
- Unused for 30 days → Demote to WARM
- Unused for 90 days → Archive to COLD
- Never automatically deleted

---

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SYSTEM.md (reported line 204)May include surrounding context.

md
- Applied 3 times/7 days → Elevate to HOT
- Unused for 30 days → Demote to WARM
- Unused for 90 days → Archive to COLD
- Never automatically deleted

---

Session Persistence

Medium
Category
Rogue Agent
Confidence
80% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SYSTEM.md (reported line 269)May include surrounding context.

3. Check Cron suggestions

cat proposals/PENDING.md

4. Add Cron Task to OpenClaw configuration

text

---

Static analysis

No suspicious patterns detected.