Back to skill

Security audit

Self-Improving Agent (Anti-Loop Hardened)

Security checks for vulnerabilities and agentic risk

Overview

This skill is purpose-related but needs review because it can persistently influence future agent behavior and its optional hooks are broader than the main skill description suggests.

Install only if you want an agent-memory workflow that writes local learning records and may affect future sessions. Prefer project-scoped hooks, avoid the global UserPromptSubmit examples, review every learning before promotion, redact sensitive details, and replace the README rm -rf install command with a backup or staged install flow pinned to a reviewed version.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (4)

T01 · Skill Instruction Hijacking

Error
Location
hooks/openclaw/handler.js:8
Finding

Bootstrap Instruction Injection Enables Persistent Agent Behavior Modification

Content
View full analysis
{ // Safety checks for event structure if (!event || typeof event !== 'object') { return; } // Only handle agent:bootstrap events if (event.type !== 'agent' || event.action !== 'bootstrap') { return; } // Safety check for context if (!event.context || typeof event.context !== 'object') { return; } // Inject the reminder as a virtual bootstrap file // Check that bootstrapFiles is an array before pushing if (Array.isArray(event.context.bootstrapFiles)) { event.context.bootstrapFiles.push({ path: 'SELF_IMPROVEMENT_REMINDER.md', content: REMINDER_CONTENT, virtual: true, }); } }; ``` The related promotion guidance states: ```markdown | Learning Type | Promote To | |---------------|------------| | Behavioral patterns | `SOUL.md` | | Workflow improvements | `AGENTS.md` | | Tool gotchas | `TOOLS.md` | ``` ### Techn ...[truncated 2612 chars]
Remediation
View remediation

T01 · Skill Instruction Hijacking

Error
Location
scripts/activator.sh:2
Finding

Global Per-Prompt Hook Injects Instructions Outside the Skill's Declared Trigger Boundaries

Content
View full analysis
After completing this task, evaluate if extractable knowledge emerged: - Non-obvious solution discovered through investigation? - Workaround for unexpected behavior? - Project-specific pattern learned? - Error required debugging to resolve? If yes: Log to .learnings/ using the self-improvement skill format. If high-value (recurring, broadly applicable): Consider skill extraction. EOF ``` The documented hook uses an empty matcher: ```json { "hooks": { "UserPromptSubmit": [ { "matcher": "", "hooks": [ { "type": "command", "command": "./skills/self-improvement/scripts/activator.sh" } ] } ] } } ``` ### Technical Analysis The empty `UserPromptSubmit` matcher causes the script to run after every user prompt. Its output is explicitly intended to become agent context and tells the agent to reinterpret task activity as material for persistent learning or Skill extraction. This behavior is broader than the trigger boundaries declared in `SKILL.md`, which state that ordinary conversations, design discussions, document cleanup, and non-agent corrections must not trigger learning capture. The hook has no programmatic mechanism to enforce those exclusions. Instead, it injects broader criteria such as any project-specific pattern or workaround. Because the instruction is inserted for unrelated prompts, an attacker does not need to ...[truncated 1445 chars]
Remediation
View remediation

T03 · Remote Payload Retrieval and Execution

Error
Location
README.md:35
Finding

Installation Fetches Mutable Remote Content After Destructively Removing the Existing Skill

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/extract-skill.sh:92
Finding

Output-Directory Validation Can Be Bypassed Through Symbolic Links

Content
View full analysis
"$SKILL_PATH/SKILL.md" << TEMPLATE --- name: $SKILL_NAME description: "[TODO: Add a concise description of what this skill does and when to use it]" --- # $(echo "$SKILL_NAME" | sed 's/-/ /g' | awk '{for(i=1;i<=NF;i++) $i=toupper(substr($i,1,1)) tolower(substr($i,2))}1') [TODO: Brief introduction explaining the skill's purpose] ## Quick Reference | Situation | Action | |-----------|--------| | [Trigger condition] | [What to do] | ## Usage [TODO: Detailed usage instructions] ## Examples [TODO: Add concrete examples] ## Source Learning This skill was extracted from a learning entry. - Learning ID: [TODO: Add original learning ID] - Original File: .learnings/LEARNINGS.md TEMPLATE ``` ### Technical Analysis The script attempts to confine output to the current workspace by rejecting absolute paths and literal `..` components. These are lexical checks only. It does not resolve the canonical path of the current directory or destination and doe ...[truncated 1777 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (32)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
96% confidence
Finding

The README includes a forceful recursive deletion command against a user-local skill directory without validation or warning. In a skill-distribution context, users are likely to copy-paste installation commands directly; that makes destructive shell guidance more dangerous because small path mistakes, symlink surprises, or environment differences can delete unintended content.

Content

Scanner excerpt · README.md (reported line 38)May include surrounding context.

bash
# 替换原版
rm -rf ~/.openclaw/skills/self-improving-agent
git clone https://github.com/lanyasheng/self-improving-agent.git ~/.openclaw/skills/self-improving-agent

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
96% confidence
Finding

The README includes a forceful recursive deletion command against a user-local skill directory without validation or warning. In a skill-distribution context, users are likely to copy-paste installation commands directly; that makes destructive shell guidance more dangerous because small path mistakes, symlink surprises, or environment differences can delete unintended content.

Content

Scanner excerpt · README.md (reported line 38)May include surrounding context.

bash
# 替换原版
rm -rf ~/.openclaw/skills/self-improving-agent
git clone https://github.com/lanyasheng/self-improving-agent.git ~/.openclaw/skills/self-improving-agent

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
96% confidence
Finding

This English installation block repeats the same unsafe recursive deletion instruction. Because README installation snippets are high-trust and commonly executed verbatim, including rm -rf in setup instructions creates a realistic risk of user-inflicted data loss even if the intended target is limited.

Content

Scanner excerpt · README.md (reported line 73)May include surrounding context.

bash
# Replace original
rm -rf ~/.openclaw/skills/self-improving-agent
git clone https://github.com/lanyasheng/self-improving-agent.git ~/.openclaw/skills/self-improving-agent

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
96% confidence
Finding

This English installation block repeats the same unsafe recursive deletion instruction. Because README installation snippets are high-trust and commonly executed verbatim, including rm -rf in setup instructions creates a realistic risk of user-inflicted data loss even if the intended target is limited.

Content

Scanner excerpt · README.md (reported line 73)May include surrounding context.

bash
# Replace original
rm -rf ~/.openclaw/skills/self-improving-agent
git clone https://github.com/lanyasheng/self-improving-agent.git ~/.openclaw/skills/self-improving-agent

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents a narrowly scoped skill for recording learnings in response to specific events such as failures, direct user corrections, missing capabilities, or tool/API issues. By contrast, the code does not detect or handle those situations directly. Instead, it always activates during agent bootstrap and injects a reminder file that broadly instructs the agent to self-evaluate after tasks. This is a materially different trigger and behavior. The injected reminder also expands scope by suggesting promotion of patterns into SOUL.md, AGENTS.md, and TOOLS.md, which goes beyond simple learning capture. Finally, the description contains important operational constraints—only one learning log per user message and no chaining—but the code does not implement or enforce them. Therefore the description does not accurately represent the actual behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description presents a reactive skill for capturing learnings/errors under specific situations, with behavioral constraints around when and how often it should act. The actual code instead runs only on agent bootstrap and injects a persistent reminder file into the agent context. This is a materially different mechanism and trigger from the declared behavior. It also expands scope by suggesting promotion of patterns to additional files beyond simple learning/error capture. Because the code neither performs nor enforces the described conditional logging behavior, the description does not accurately represent the implementation.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents a narrowly scoped learning-capture skill intended for specific situations such as unexpected failures, user corrections, missing capabilities, or external tool/API failures, with explicit operational constraints. The actual code is an always-on activator hook for UserPromptSubmit that inserts a generic reminder after every prompt, asking the agent to assess whether any useful knowledge emerged from the task. This is materially broader in trigger and scope than declared, because it is not limited to the listed conditions. It also introduces an additional behavior—considering skill extraction for high-value learnings—not mentioned in the description. Finally, the code does not enforce the stated safeguards of one learning log per message and no chaining. These differences are substantial enough to count as a mismatch rather than mere implementation detail.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The code only covers one narrow mechanism: detecting likely command failures from Bash tool output and printing a reminder. That partially aligns with the declared purpose around command/tool failures, but the description presents a broader self-improvement skill spanning multiple trigger conditions and behavioral constraints. The supplied code does not handle user corrections or capability-gap requests, and it does not enforce the stated 'maximum 1 learning log per user message' or 'do NOT chain' constraints. It also introduces an undeclared operational detail: recommending logging to a specific file path with a specific identifier format. Because the actual behavior is materially narrower in scope while also adding an undeclared resource/path-specific behavior, this is a description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description is about recording learnings/corrections in response to failures, user corrections, missing capabilities, or tool/API errors. The actual code does not capture learnings, log errors for continuous improvement, enforce one-learning-per-message limits, or react to the declared conversational triggers. Instead, it is a standalone shell script whose primary function is to generate a new skill directory and SKILL.md template file from a skill name, optionally in dry-run mode. This is a materially different purpose and introduces undeclared filesystem write capabilities. The script’s references to 'learning entry' are only in comments/template text and do not implement the declared learning-log behavior.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Directing users to modify ~/.claude/settings.json introduces persistent agent configuration changes at the user scope. In this context, that is security-relevant because it grants durable hook execution across all future sessions, increasing blast radius if the configured script is unsafe, replaced, or later modified.

Content

Scanner excerpt · references/hooks-setup.md (reported line 48)May include surrounding context.

Option 2: User-Level Configuration

Add to ~/.claude/settings.json for global activation:

json
{

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The security section claims the scripts only output text and do not run commands, but the documented integration explicitly executes shell scripts as hook commands. This is dangerous because it downplays execution risk, may cause operators to trust active code paths too broadly, and can lead to unsafe deployment of command-executing hooks with the agent's privileges.

Content

No source excerpt is available for this finding.

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

The sessions_send example demonstrates transmitting learning content to another session, which is an information transfer primitive and can be used to move potentially sensitive data outside the originating context. In a self-improvement skill, this is especially risky because the source material may include error details, user corrections, or tool outputs that were never intended for cross-session disclosure.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 177)May include surrounding context.

sessions_send

Send message to another session:

text
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The OpenClaw-specific trigger table can cause multiple persistent writes from one user message, contradicting the manifest requirement of at most one learning log per user message. Violating that constraint can lead to runaway persistence, over-collection, and larger attack surface for prompt poisoning or accidental state corruption.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The README instructs users to run a destructive rm -rf command to remove an existing skill directory, but provides no warning, verification step, or safer alternative. While the path is scoped to a specific directory rather than arbitrary user data, documentation that normalizes forceful recursive deletion increases the chance of accidental data loss if the path is mistyped, expanded unexpectedly, or copied into the wrong environment.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill instructs the agent to append to persistent workspace files automatically, but it provides no requirement to notify the user that local state will be modified or retained. In an agent environment, silent persistence can surprise users, alter workspace integrity, and create audit/privacy concerns, especially when triggered by routine corrections or failures.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill directs the agent to persist corrections, errors, and feature requests into local memory files, which can capture sensitive operational details and user-provided context for later resurfacing. Because this is a self-improvement skill used across conversations, the persistence context makes retention more dangerous: routine interactions may accumulate private data without clear minimization or deletion controls.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The feature-request template explicitly asks for 'What the user wanted to do' and 'Why they needed it', encouraging storage of potentially sensitive business, personal, or security-relevant context in durable files. This increases the risk of privacy leakage and later prompt-surface exposure because future sessions may read and reuse those details outside the original context.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest context says this skill should be used only in response to failures, explicit user corrections, missing capabilities, or external tool/API failures, with at most one learning log per user message. This hook instead runs unconditionally on agent:bootstrap and injects a general reminder to check .learnings/ and log corrections, errors, and discoveries, which broadens the behavior beyond the stated event-driven self-improvement use cases.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · references/examples.md (reported line 301)May include surrounding context.

When the above learning is extracted as a skill, it becomes:

File: skills/docker-m1-fixes/SKILL.md

markdown
---

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The guide instructs users to trigger self-improvement evaluation after every prompt, which exceeds the skill's stated narrow trigger conditions. This creates unnecessary always-on behavior, increasing prompt-surface exposure and the chance of persistent logging, instruction drift, or unintended capture of sensitive context.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
83% confidence
Finding

The guide creates persistent session-affecting configuration in .claude/settings.json, causing the self-improvement behavior to survive beyond a single interaction. This is risky in an agent skill because persistent reminders and hook execution can silently alter future agent behavior and expand the attack surface over time.

Content

Scanner excerpt · references/hooks-setup.md (reported line 15)May include surrounding context.

Option 1: Project-Level Configuration

Create .claude/settings.json in your project root:

json
{

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

An empty matcher causes the hook to fire on every user prompt, creating an overly broad activation scope. In this skill context, always-on execution is more dangerous because the feature is meant for limited self-improvement events, yet it becomes a persistent interceptor that can amplify logging and prompt-context handling unnecessarily.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The user-level configuration enables global always-on invocation through an unconstrained matcher in a persistent home-directory settings file. This magnifies risk because the behavior applies across projects and sessions, potentially causing unintended execution and collection in unrelated contexts.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The Codex CLI example also uses an empty matcher, leaving the trigger effectively always-on and unspecified. Because this is documentation users are likely to copy verbatim, it propagates a broad execution pattern into another agent environment without clear need or scope control.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 57)May include surrounding context.

openclaw hooks enable self-improvement

text

### 3. Create Learning Files

Create the `.learnings/` directory in your workspace:

Static analysis

Detected: suspicious.destructive_delete_command

Documentation contains a destructive delete command without an explicit confirmation gate.

Warn
Code
suspicious.destructive_delete_command
Location
README.md:38