Back to skill

Security audit

Self Improvement Local

Security checks for vulnerabilities and agentic risk

Overview

The skill has a clear self-improvement purpose, but it broadly persists conversation-derived guidance into future agent behavior and optional hooks without enough scoping or approval controls.

Install only if you want persistent local learning files and possible future changes to agent instruction files. Keep hooks project-scoped, avoid global empty-match hooks, review every promotion into CLAUDE.md/AGENTS.md/SOUL.md/TOOLS.md, and do not log secrets or raw transcripts.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:348
Finding

Conversation-Derived Content Can Be Promoted into Persistent Agent Instructions

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:348-380, SKILL.md:462-468, references/openclaw-integration.md:72-108
Vulnerability Type: Persistent agent memory poisoning
Risk Level: High

Vulnerable Code

markdown
### Promotion Rule (System Prompt Feedback)

Promote recurring patterns into agent context/system prompt files when all are true:

- `Recurrence-Count >= 3`
- Seen across at least 2 distinct tasks
- Occurred within a 30-day window

Promotion targets:
- `CLAUDE.md`
- `AGENTS.md`
- `.github/copilot-instructions.md`
- `SOUL.md` / `TOOLS.md` for OpenClaw workspace-level guidance when applicable

Write promoted rules as short prevention rules (what to do before/while coding),
not long incident write-ups.

The general promotion guidance is even broader:

markdown
## Best Practices

1. **Log immediately** - context is freshest right after the issue
2. **Be specific** - future agents need to understand quickly
3. **Include reproduction steps** - especially for errors
4. **Link related files** - makes fixes easier
5. **Suggest concrete fixes** - not just "investigate"
6. **Use consistent categories** - enables filtering
7. **Promote aggressively** - if in doubt, add to CLAUDE.md or .github/copilot-instructions.md
8. **Review regularly** - stale learnings lose value

The OpenClaw integration also directs learnings into automatically loaded workspace files:

markdown
### Promotion Decision Tree

Is the learning project-specific?
├── Yes → Keep in .learnings/
└── No → Is it behavioral/style-related?
    ├── Yes → Promote to SOUL.md
    └── No → Is it tool-related?
        ├── Yes → Promote to TOOLS.md
        └── No → Promote to AGENTS.md (workflow)

Technical Analysis

The Skill records information originating from conversations, including user corrections, behavioral observations, and discovered patterns. It then recommends conve ...[truncated 2072 chars]

Remediation
View remediation

Remediation Suggestions

  1. Require explicit, informed user approval immediately before modifying any persistent agent instruction file.
  2. Treat all conversation-derived material as untrusted data, including repeated corrections.
  3. Never promote user-provided text verbatim. Convert it into a proposed rule and present the exact diff for approval.
  4. Store provenance with every proposal, including source session, author, creation time, supporting evidence, and affected scope.
  5. Add a mandatory security review that rejects rules which alter safety boundaries, authorization behavior, secret handling, external communication, or tool permissions.
  6. Replace “promote aggressively” with a conservative, evidence-based process.
  7. Keep learning records in a non-instruction data file by default. Do not automatically load raw records as agent instructions.
  8. Introduce expiration and rollback mechanisms for promoted rules.
  9. Restrict writes to SOUL.md, AGENTS.md, TOOLS.md, and similar files through an allowlisted update function with audit logging.
  10. Require independent evidence rather than recurrence count alone before accepting a rule as valid.

T01 · Skill Instruction Hijacking

Warning
Location
scripts/activator.sh:8
Finding

Broad Hooks Inject Skill-Controlled Instructions into Unrelated Agent Sessions

Content
View full analysis

Vulnerability Details

File Location: scripts/activator.sh:8-19, hooks/openclaw/handler.js:8-25, hooks/openclaw/handler.js:44-52, SKILL.md:491-536
Vulnerability Type: Global prompt and bootstrap instruction injection
Risk Level: Medium

Vulnerable Code

bash
# Output reminder as system context
cat << 'EOF'
<self-improvement-reminder>
After completing this task, evaluate if extractable knowledge emerged:
- Non-obvious solution discovered through investigation?
- Workaround for unexpected behavior?
- Project-specific pattern learned?
- Error required debugging to resolve?

If yes: Log to .learnings/ using the self-improvement skill format.
If high-value (recurring, broadly applicable): Consider skill extraction.
</self-improvement-reminder>
EOF

The OpenClaw bootstrap hook injects similar instructions as a virtual prompt file:

javascript
const REMINDER_CONTENT = `
## Self-Improvement Reminder

After completing tasks, evaluate if any learnings should be captured:

**Log when:**
- User corrects you → \`.learnings/LEARNINGS.md\`
- Command/operation fails → \`.learnings/ERRORS.md\`
- User wants missing capability → \`.learnings/FEATURE_REQUESTS.md\`
- You discover your knowledge was wrong → \`.learnings/LEARNINGS.md\`
- You find a better approach → \`.learnings/LEARNINGS.md\`

**Promote when pattern is proven:**
- Behavioral patterns → \`SOUL.md\`
- Workflow improvements → \`AGENTS.md\`
- Tool gotchas → \`TOOLS.md\`

Keep entries simple: date, title, what happened, what to do differently.
`.trim();
javascript
if (Array.isArray(event.context.bootstrapFiles)) {
  event.context.bootstrapFiles.push({
    path: 'SELF_IMPROVEMENT_REMINDER.md',
    content: REMINDER_CONTENT,
    virtual: true,
  });
}

Technical Analysis

Once enabled, the UserPromptSubmit hook runs with an empty matcher, meaning it applies to ...[truncated 1697 chars]

Remediation
View remediation

Remediation Suggestions

  1. Replace the empty matcher with narrow, task-specific matchers.
  2. Avoid recommending user-level global activation by default.
  3. Make logging a separate user-approved action rather than an automatic secondary objective.
  4. Require confirmation before each filesystem write or persistent promotion.
  5. Add explicit instructions that higher-priority policies, user privacy requirements, and the current task always take precedence.
  6. Disable promotion instructions in hook-generated context by default.
  7. Add technical redaction for tokens, credentials, environment variables, private keys, and common secret formats.
  8. Provide a session-only mode that produces no persistent state.
  9. Clearly display the exact scope and affected repositories before global activation.
  10. Add automated tests confirming that unrelated tasks do not trigger persistent writes.

T09 · Insecure Skill Coding Practices

Warning
Location
hooks/openclaw/handler.js:27
Finding

Executed JavaScript Hook Omits the TypeScript Sub-Agent Exclusion

Content
View full analysis

Vulnerability Details

File Location: hooks/openclaw/handler.js:27-52, hooks/openclaw/handler.ts:43-61
Vulnerability Type: Divergent security behavior between source and distributed implementation
Risk Level: Medium

Vulnerable Code

The TypeScript implementation contains a sub-agent exclusion:

typescript
// Skip sub-agent sessions to avoid bootstrap issues
// Sub-agents have sessionKey patterns like "agent:main:subagent:..."
const sessionKey = event.sessionKey || '';
if (sessionKey.includes(':subagent:')) {
  return;
}

// Inject the reminder as a virtual bootstrap file
// Check that bootstrapFiles is an array before pushing
if (Array.isArray(event.context.bootstrapFiles)) {
  event.context.bootstrapFiles.push({
    path: 'SELF_IMPROVEMENT_REMINDER.md',
    content: REMINDER_CONTENT,
    virtual: true,
  });
}

The JavaScript implementation proceeds directly from context validation to injection:

javascript
// Safety check for context
if (!event.context || typeof event.context !== 'object') {
  return;
}

// Inject the reminder as a virtual bootstrap file
// Check that bootstrapFiles is an array before pushing
if (Array.isArray(event.context.bootstrapFiles)) {
  event.context.bootstrapFiles.push({
    path: 'SELF_IMPROVEMENT_REMINDER.md',
    content: REMINDER_CONTENT,
    virtual: true,
  });
}

Technical Analysis

The TypeScript source explicitly excludes session keys containing :subagent:. The distributed JavaScript handler does not implement this safeguard.

If OpenClaw loads handler.js, sub-agent bootstrap events pass the event and context checks and receive the injected self-improvement reminder. This makes actual runtime behavior broader than the TypeScript source suggests.

Maintaining manually divergent source and runtime files also weakens code review: reviewers may approve the TypeScript safeguard while the executable JavaS ...[truncated 956 chars]

Remediation
View remediation

Remediation Suggestions

  1. Regenerate handler.js from the corrected TypeScript source.
  2. Use one authoritative source file and a reproducible build process.
  3. Fail the build when generated JavaScript differs from compiled TypeScript.
  4. Add tests covering main-agent and sub-agent session keys.
  5. Verify which entry point OpenClaw actually loads and document it explicitly.
  6. Prefer an explicit event property identifying sub-agents rather than relying only on a session-key substring.
  7. Add package integrity checks so stale generated files cannot be released.

T08 · Insecure Dependencies

Warning
Location
SKILL.md:50
Finding

Installation Instructions Use Mutable, Unpinned Upstream Sources

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:50-59, references/openclaw-integration.md:31-41
Vulnerability Type: Unpinned Skill supply-chain installation
Risk Level: Medium

Vulnerable Code

markdown
**Via ClawdHub (recommended):**
```bash
clawdhub install self-improving-agent

Manual:

bash
git clone https://github.com/peterskoett/self-improving-agent.git ~/.openclaw/skills/self-improving-agent
text

The integration guide repeats the mutable installation method:

```markdown
### 1. Install the Skill

```bash
clawdhub install self-improving-agent

Or copy manually:

bash
cp -r self-improving-agent ~/.openclaw/skills/
text

### Technical Analysis

The registry installation command does not specify the audited version, and the Git command clones the repository's current default branch without an immutable commit reference or integrity verification.

The destination is an automatically loaded Skill directory. Therefore, a compromised registry release, owner account, repository, or default branch could introduce altered instructions or executable hooks after the reviewed version.

The Git command is not piped into a shell, which limits immediate execution risk. The supply-chain concern arises when the installed Skill is later loaded or its hooks are enabled.

### Attack Path

1. An attacker compromises the upstream registry package, publisher account, repository, or default branch.
2. The attacker publishes modified Skill instructions or malicious hook scripts.
3. A user follows the unpinned installation command.
4. The mutable upstream content is installed under the active OpenClaw Skill directory.
5. OpenClaw loads the altered Skill instructions.
6. If the user enables the hook, attacker-controlled code executes with the agent process's permissions on matching events.

### Impact Assessment

Successful exploitation would provide the malicio
...[truncated 455 chars]
Remediation
View remediation

Remediation Suggestions

  1. Pin registry installation instructions to the audited version, such as version 3.0.13, where supported.
  2. Pin manual Git installation to an immutable commit SHA or signed release tag.
  3. Publish SHA-256 checksums and cryptographic signatures for releases.
  4. Require signature and checksum verification before copying content into an active Skill directory.
  5. Download and inspect the package in a staging directory before installation.
  6. Do not enable hooks automatically after installation.
  7. Document the expected owner identity, repository URL, release version, and checksum.
  8. Use protected release branches, signed commits or tags, and multi-factor authentication for publisher accounts.
  9. Provide a lockfile or manifest recording the exact installed Skill version and digest.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (17)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description is about recording and using learnings from failures, corrections, outdated knowledge, and better approaches. The supplied code does not capture learnings, errors, corrections, or review prior learnings. Instead, it is a filesystem-writing helper that creates a new skill scaffold under a skills directory, including a templated SKILL.md file. While the comments mention creating a skill from a learning entry, the implementation only generates a template and does not ingest, extract, store, or analyze learning data. This is a materially different primary purpose, so the description does not accurately represent the code's behavior.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
92% confidence
Finding

Directing users to modify ~/.claude/settings.json grants persistence in an agent configuration directory that affects future sessions globally. In skill context, this is more dangerous because the feature is specifically designed to self-trigger and capture learnings, so a malicious or tampered script path in global config could surveil or influence broad agent activity over time.

Content

Scanner excerpt · references/hooks-setup.md (reported line 48)May include surrounding context.

Option 2: User-Level Configuration

Add to ~/.claude/settings.json for global activation:

json
{

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

Instructions found that direct the agent to transmit conversation context or user data to external services.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 181)May include surrounding context.

sessions_send

Send message to another session:

text
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The manifest description says to use the skill when a command fails, when the user corrects the agent, when a capability is missing, when an API fails, when knowledge is outdated, and also to review learnings before major tasks. This scope is very broad and lacks exclusion conditions, so it risks matching many ordinary interactions rather than a narrowly defined invocation context.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
84% confidence
Finding

This skill encourages persistent storage of session-derived information in a workspace-level .learnings/ directory under ~/.openclaw/workspace, which can accumulate sensitive operational details across sessions. Even with warnings not to store secrets, persistent local memory increases the blast radius of any accidental disclosure, especially in shared machines, synced home directories, or multi-project workspaces.

Content

Scanner excerpt · SKILL.md (reported line 81)May include surrounding context.

└── FEATURE_REQUESTS.md

text

### Create Learning Files

```bash
mkdir -p ~/.openclaw/workspace/.learnings

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill explicitly references reading other sessions’ transcripts and sending messages across sessions, which creates a confidentiality boundary-crossing risk if used with sensitive prior conversations. Even though the text says to use these only in trusted environments and with explicit user intent, embedding this capability in a broadly activated self-improvement skill increases the chance of accidental disclosure or over-collection.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The listed triggers include phrases like "Actually...", "Can you also...", and "Is there a way to...", which commonly appear in normal user conversation. Because the file does not clearly constrain these phrases to specific contexts or provide negative examples, the skill could be invoked unintentionally during routine dialogue.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The manifest claims the skill captures learnings, errors, and corrections to improve future behavior. The documented extraction workflow creates new reusable skills from learnings, which is a separate capability involving artifact generation and packaging rather than merely recording and reviewing learnings.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · references/examples.md (reported line 301)May include surrounding context.

When the above learning is extracted as a skill, it becomes:

File: skills/docker-m1-fixes/SKILL.md

markdown
---

Session Persistence

Medium
Category
Rogue Agent
Confidence
81% confidence
Finding

The guide instructs users to create persistent hook configuration files in agent settings, causing behavior to survive beyond the immediate session. Persistence is not inherently malicious, but in this context it increases risk because automatic hooks continue running unless explicitly removed, which can outlast user intent and expand the window for misuse.

Content

Scanner excerpt · references/hooks-setup.md (reported line 15)May include surrounding context.

Option 1: Project-Level Configuration

Create .claude/settings.json in your project root:

json
{

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Using an empty matcher causes the UserPromptSubmit hook to fire on every prompt, creating a broad automatic execution surface. In the context of self-improvement hooks, this means local scripts run continuously across all interactions, increasing exposure to prompt-derived sensitive data, accidental behavior, and abuse if the script is modified or replaced.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The user-level example installs the hook into ~/.claude/settings.json with no trigger constraints, enabling automatic execution across all future sessions and repositories. Persistent, global activation materially raises risk because a compromised or overly permissive script can observe or influence every interaction, not just one project.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Although described as minimal, this setup still uses an empty matcher that executes on every prompt. Reducing the number of hooks lowers overhead, but it does not address the core issue of unconstrained automatic triggering of a command hook.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The Codex example repeats the same empty-matcher pattern, extending broad automatic command execution to another agent environment without scope limits. Reproducing unsafe defaults across tools amplifies the chance of widespread over-permissioned deployment.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The document’s security section understates risk by claiming the hook scripts only output text and do not run commands, even though they are explicitly configured as command hooks and the guide also references an extraction script that performs operational actions. This kind of misleading assurance can cause users to grant trust or enable hooks without understanding that arbitrary local scripts will execute in their agent environment.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 57)May include surrounding context.

openclaw hooks enable self-improvement

text

### 3. Create Learning Files

Create the `.learnings/` directory in your workspace:

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This markdown file includes trigger definitions, and 'Knowledge gaps' does not specify what concrete condition qualifies or when the skill should not act. Because it is a common, subjective situation in normal interactions, it can overlap with routine usage and lead to overly broad invocation behavior.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.