Back to skill

Security audit

Approved Self Improvement

Security checks for vulnerabilities and agentic risk

Overview

The skill is a mostly disclosed self-improvement logger, but it can persistently change future agent behavior and auto-update skills using weak approval records.

Install only if you are comfortable with a self-improvement workflow that writes persistent local learning logs and may affect future agent behavior. Keep auto-update disabled, require explicit approval with diffs before any skill or agent-instruction-file changes, avoid global hooks, and do not use cross-session transcript or messaging features in untrusted workspaces.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:293
Finding

Unreviewed Learnings Can Be Promoted into Persistent Agent Instructions

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:293-320
Vulnerability Type: Persistent agent memory and instruction poisoning
Risk Level: High

Vulnerable Code

markdown
## Promoting to Project Memory

When a learning is broadly applicable (not a one-off fix), promote it to permanent project memory.

### When to Promote

- Learning applies across multiple files/features
- Knowledge any contributor (human or AI) should know
- Prevents recurring mistakes
- Documents project-specific conventions

### Promotion Targets

| Target | What Belongs There |
|--------|-------------------|
| `CLAUDE.md` | Project facts, conventions, gotchas for all Claude interactions |
| `AGENTS.md` | Agent-specific workflows, tool usage patterns, automation rules |
| `.github/copilot-instructions.md` | Project context and conventions for GitHub Copilot |
| `SOUL.md` | Behavioral guidelines, communication style, principles (OpenClaw workspace) |
| `TOOLS.md` | Tool capabilities, usage patterns, integration gotchas (OpenClaw workspace) |

### How to Promote

1. **Distill** the learning into a concise rule or fact
2. **Add** to appropriate section in target file (create file if needed)
3. **Update** original entry:
   - Change `**Status**: pending` → `**Status**: promoted`
   - Add `**Promoted**: CLAUDE.md`, `AGENTS.md`, or `.github/copilot-instructions.md`

Additional reinforcement appears in SKILL.md:377-391, where recurring patterns are promoted into agent context or system-prompt files, and in SKILL.md:499, which states:

markdown
7. **Promote aggressively** - if in doubt, add to CLAUDE.md or .github/copilot-instructions.md

The optional bootstrap hook also reinforces promotion into persistent instruction files in hooks/openclaw/handler.js:30-33:

javascript
**Promote when pattern is proven:**
- Behavioral patterns → \`SOUL.md\`
- Workflow improvements → \`AGENTS.md\`
- Too
...[truncated 2596 chars]
Remediation
View remediation

Remediation Suggestions

  1. Require explicit user confirmation before every write to an agent instruction or persistent-context file.
  2. Present the exact destination, proposed text, and unified diff before requesting approval.
  3. Never promote content copied directly from command output, repository documents, external APIs, issue text, or transcripts.
  4. Track provenance and trust level for every learning. Permit promotion only from trusted, user-verified sources.
  5. Store ordinary learnings in a dedicated data-only file that is not interpreted as agent instructions.
  6. Treat CLAUDE.md, AGENTS.md, .github/copilot-instructions.md, SOUL.md, and TOOLS.md as security-sensitive configuration.
  7. Add validation that rejects instructions involving credential access, safety-constraint changes, hidden execution, unapproved network access, or privilege expansion.
  8. Replace “promote aggressively” with a default-deny policy requiring deliberate review.
  9. Record an immutable audit entry containing the approval event, source learning, exact diff, timestamp, and destination.
  10. Provide a rollback mechanism for every promoted rule.

T05 · Unauthorized Access and Privilege Escalation

Error
Location
SKILL.md:702
Finding

Mutable Markdown File Is Trusted as Authorization for Autonomous Skill Modification

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:702-751
Vulnerability Type: Forgeable authorization and unsafe access-control decision
Risk Level: High

Vulnerable Code

markdown
### Auto-Update Authorization

By default, **no skill is authorized for auto-update**. The user must explicitly grant permission on a per-skill basis.

**How to authorize a skill for auto-update:**

The user says something like:
- "Allow auto-updates for the [skill-name] skill"
- "The [skill-name] skill can self-improve without asking me"
- "Auto-approve improvements to [skill-name]"

When authorized:

1. Add an entry to `.learnings/AUTO_UPDATE_AUTHORIZATIONS.md`:

```markdown
## [skill-name]
**Authorized**: ISO-8601 timestamp
**Authorized By**: user
**Scope**: full
**Notes**: User authorized auto-update in conversation

---
  1. Confirm to the user: "I've authorized auto-updates for [skill-name]. Future improvements to this skill will be applied automatically and I'll inform you what changed. You can revoke this anytime."

How to revoke auto-update:

The user says something like:

  • "Stop auto-updating [skill-name]"
  • "Require approval for [skill-name] again"
  • "Revoke auto-update for [skill-name]"

When revoked: Remove the entry from AUTO_UPDATE_AUTHORIZATIONS.md and confirm.

Checking authorization before applying changes:

bash
grep -l "## skill-name-here" .learnings/AUTO_UPDATE_AUTHORIZATIONS.md 2>/dev/null

If the skill is listed AND its entry is present (not just the file header), it is authorized. Otherwise, require user approval.

Auto-Update Behavior

When a skill IS authorized for auto-update and a failure is detected:

  1. Analyze the failure and determine the fix
  2. Check if a pending proposal already exists — if so, apply it
  3. If no proposal exists, create one with status applied (for the record) and apply the fix
  4. Always inform the user wh ...[truncated 2689 chars]
Remediation
View remediation

Remediation Suggestions

  1. Do not use a workspace Markdown file as the authoritative permission store.
  2. Store authorization in a platform-managed, access-controlled configuration system outside the repository and ordinary agent-writable paths.
  3. Require fresh user confirmation for every Skill modification, even where a standing preference exists.
  4. Bind approval to the exact Skill identifier, canonical path, requested files, permitted change types, and final content hash.
  5. Display the complete proposed diff before approval and reject changes that differ from the approved hash.
  6. If persistent authorization is unavoidable, use a structured format with strict parsing, schema validation, expiry, explicit scope, revocation state, and integrity protection.
  7. Enforce minor_only or other scopes in code rather than documenting them without validation.
  8. Reject authorizations found in cloned, untrusted, or repository-controlled files.
  9. Prevent automatic modification of executable scripts under any standing authorization; require per-change approval for code.
  10. Maintain an append-only audit trail showing the authenticated approval source, timestamp, requested change, resulting diff, and rollback information.

T08 · Insecure Dependencies

Warning
Location
SKILL.md:56
Finding

Installation Instructions Retrieve Mutable and Unpinned Skill Content

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:56-68
Vulnerability Type: Unpinned external supply-chain installation
Risk Level: Medium

Vulnerable Code

markdown
### Installation

**Via ClawdHub (recommended):**
```bash
clawdhub install self-improving-agent

Manual:

bash
git clone https://github.com/peterskoett/self-improving-agent.git ~/.openclaw/skills/self-improving-agent

Remade for openclaw from original repo : https://github.com/pskoett/pskoett-ai-skills - https://github.com/pskoett/pskoett-ai-skills/tree/main/skills/self-improvement

text

### Technical Analysis

Both installation methods retrieve the latest mutable upstream content without pinning a reviewed release, commit hash, checksum, or cryptographic signature. The manual installation source is also different from the repositories identified as the original source in the attribution statement.

Agent Skill documents and hooks are security-sensitive even when they are not conventional compiled dependencies. Skill instructions influence agent behavior, and hook JavaScript or shell scripts may execute with the agent's permissions. Therefore, installing mutable content after an audit allows the effective code and instruction payload to change without further review.

This finding does not establish that the current upstream repository is malicious. The vulnerability is the absence of version and integrity controls, which creates a supply-chain attack opportunity.

### Attack Path

1. An attacker compromises an upstream repository, package-publishing account, distribution channel, maintainer account, or repository ownership.
2. The attacker changes the Skill instructions or bundled hook scripts after this version has been reviewed.
3. A user follows the documented `clawdhub install` or `git clone` command.
4. The installer retrieves the current attacker-modified version because no immutable revision is specified.
5.
...[truncated 761 chars]
Remediation
View remediation

Remediation Suggestions

  1. Pin installation to a specific reviewed release or full Git commit hash.
  2. Publish SHA-256 or stronger checksums for all distributed files and require verification during installation.
  3. Use signed releases or signed commits and document signature verification.
  4. Identify one canonical upstream repository and clearly explain the relationship to any forks or original sources.
  5. Avoid recommending installation directly from a mutable default branch.
  6. Separate Skill documents from executable hooks and require an additional explicit review before hooks are enabled.
  7. Provide a manifest containing expected file paths and hashes so unexpected additions are rejected.
  8. Re-run security review whenever the pinned revision changes.
  9. Use least-privilege execution for hooks and prevent them from accessing credentials or unrelated workspace data unless explicitly required.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (23)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description presents a skill for recording failures/corrections, reviewing learnings, tracking pending improvements, and ensuring skill modifications only occur with explicit user approval. The supplied code does something narrower and different: it listens only for agent bootstrap events and injects a virtual markdown reminder file into bootstrapFiles. This is a distinct operational capability and trigger that is not disclosed. While the reminder content is related to self-improvement, the code does not actually perform the core described functions such as capturing entries, reviewing proposals, matching failures to proposals, or handling approval workflows for skill updates. Therefore the description does not accurately represent the actual code behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description is about a meta-skill that records failures/corrections, recommends or reviews skill improvements, and enforces user approval before modifying skills. The supplied code does not implement that behavior. Instead, it is a file-generation helper that scaffolds a new skill directory and template markdown file. While the comments mention creating a skill from a learning entry, the script neither captures learnings nor reviews pending proposals nor checks for explicit user approval. Its primary purpose and concrete capabilities are materially different from the declared description, so this is a clear mismatch.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Directing users to modify ~/.claude/settings.json establishes persistent execution from the agent configuration directory, which is a sensitive trust boundary. If the referenced script path is later altered, replaced, or pointed at unreviewed content, the hook will continue running automatically in future sessions.

Content

Scanner excerpt · references/hooks-setup.md (reported line 48)May include surrounding context.

Option 2: User-Level Configuration

Add to ~/.claude/settings.json for global activation:

json
{

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

Instructions found that direct the agent to transmit conversation context or user data to external services.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 181)May include surrounding context.

sessions_send

Send message to another session:

text
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")

Session Persistence

Medium
Category
Rogue Agent
Confidence
78% confidence
Finding

Writing persistent data under ~/.openclaw/workspace/.learnings creates session persistence beyond the current task, which can accumulate sensitive operational context over time. While the skill advises against logging secrets, persistence in a shared or synced workspace still raises confidentiality and retention concerns.

Content

Scanner excerpt · SKILL.md (reported line 91)May include surrounding context.

└── IMP-YYYYMMDD-XXX-skill-name.md

text

### Create Learning Files

```bash
mkdir -p ~/.openclaw/workspace/.learnings/pending-improvements

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The trigger phrase 'What needs updating?' is broad enough to match ordinary maintenance conversations, which can cause the skill to activate in contexts the user did not intend. Ambiguous activation in an agentic system can lead to unsolicited review of pending proposals or preparation for skill changes, increasing the chance of accidental state changes or persuasive overreach.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
87% confidence
Finding

Authorizing a skill to 'self-improve without asking' enables autonomous modification of prompt artifacts, which weakens human oversight over behavior changes. In practice this can allow subtle prompt drift, persistence of bad instructions, or insertion of risky automation into future sessions under the guise of approved self-maintenance.

Content

Scanner excerpt · SKILL.md (reported line 708)May include surrounding context.

md
The user says something like:
- "Allow auto-updates for the [skill-name] skill"
- "The [skill-name] skill can self-improve without asking me"
- "Auto-approve improvements to [skill-name]"

When authorized:

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
87% confidence
Finding

The phrase 'Auto-approve improvements' explicitly permits bypassing interactive review for future changes, creating an avenue for unattended modification of agent behavior. Because skills shape future decision-making, this persistence mechanism can amplify mistakes or maliciously crafted proposals once approval boundaries are relaxed.

Content

Scanner excerpt · SKILL.md (reported line 709)May include surrounding context.

md
The user says something like:
- "Allow auto-updates for the [skill-name] skill"
- "The [skill-name] skill can self-improve without asking me"
- "Auto-approve improvements to [skill-name]"

When authorized:

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill expands from logging learnings into creating new reusable skills, which broadens its authority and can normalize writing executable prompt artifacts to disk. In an agent ecosystem, this increases the chance that unreviewed or over-broad behaviors become persisted as new skills and later auto-loaded or trusted beyond the user's original intent.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · assets/AUTO_UPDATE_AUTHORIZATIONS.md (reported line 3)May include surrounding context.

md
# Auto-Update Authorizations

Skills listed here have been explicitly authorized by the user for automatic self-improvement without approval. By default, **no skill** is authorized for auto-update. The user must explicitly grant auto-update permission on a per-skill basis.

## Format

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · references/examples.md (reported line 301)May include surrounding context.

When the above learning is extracted as a skill, it becomes:

File: skills/docker-m1-fixes/SKILL.md

markdown
---

Session Persistence

Medium
Category
Rogue Agent
Confidence
82% confidence
Finding

Placing hook configuration in .claude/settings.json introduces session-persistent behavior that automatically reactivates across future agent runs in that project. Persistence is not inherently malicious, but it is security-relevant because users may forget the hook remains active and continues executing commands with agent privileges.

Content

Scanner excerpt · references/hooks-setup.md (reported line 15)May include surrounding context.

Option 1: Project-Level Configuration

Create .claude/settings.json in your project root:

json
{

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

An empty matcher causes the hook to fire on every submitted prompt, creating an overly broad automatic execution point. Even if the current script is intended to be lightweight, broad triggering expands attack surface, increases exposure to sensitive context, and normalizes always-on command execution in agent sessions.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The user-level setup enables automatic command-hook execution across all sessions via ~/.claude/settings.json without narrowing when it runs. Global persistence magnifies the blast radius of any buggy, changed, or replaced script and can expose unrelated projects and prompts to the hook.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The Codex CLI example repeats the empty matcher pattern, again causing unconditional execution on every prompt. Reproducing this broad trigger in multiple tools increases the likelihood that users copy an unsafe default into persistent configuration.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The security section makes a materially misleading claim that the scripts only emit text and do not run commands or modify files, while the document elsewhere explicitly configures command hooks and references an extraction script that can scaffold a skill. This can cause operators to grant trust or enable hooks under false assumptions, increasing the chance of unsafe execution with the agent's privileges.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 57)May include surrounding context.

openclaw hooks enable self-improvement

text

### 3. Create Learning Files

Create the `.learnings/` directory in your workspace:

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The document instructs promoting learnings from transient notes into persistent workspace files such as SOUL.md, TOOLS.md, and AGENTS.md without clearly requiring explicit user approval for each such update. In this skill’s context, those files are injected into future sessions, so writing to them changes future model behavior and can create persistent prompt-injection or policy drift across sessions.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The standard trigger list is broad enough that routine model uncertainty, ordinary API errors, or generic command failures may invoke the self-improvement workflow too often. Over-broad activation increases the chance of capturing adversarial or low-quality inputs as learnings and can lead to unnecessary persistence or unsafe update proposals.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The OpenClaw-specific trigger table tells the agent to log tool errors, handoff confusion, and behavior surprises directly into persistent prompt files. Because these files are loaded as future session context, this effectively authorizes autonomous self-modification contrary to the manifest’s explicit approval requirement and could persist attacker-influenced content.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The OpenClaw-specific trigger descriptions specify destination files but not trust boundaries, approval requirements, or constraints on what content may be captured. This ambiguity makes it easy for benign operational events or attacker-crafted inputs to be transformed into persistent instructions that affect future sessions.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

Documenting cross-session transcript access introduces a privacy and data-exposure pathway that is not necessary for the core purpose of local learning/error logging. Even with cautionary language, encouraging transcript reads or sends across sessions can lead to unintended disclosure of sensitive conversation content or contextual data between tasks.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

The invocation phrase is overly broad and can plausibly match ordinary user conversation, causing the skill to activate outside the user's intended context. In a self-improvement skill, unintended activation is risky because it may surface pending proposals, influence workflow, or prompt update-related actions when the user was only asking a general question.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.