Back to skill

Security audit

Self Improving Agent

Security checks for vulnerabilities and agentic risk

Overview

The skill is not clearly malicious, but it can persistently change agent memory and hooks while capturing session context too broadly.

Install only if you want a persistent agent-memory workflow. Keep it project-scoped, avoid global hooks, review every hook script and settings change, never store secrets or raw transcripts in .learnings, and require a visible diff plus explicit approval before promoting anything into CLAUDE.md, AGENTS.md, SOUL.md, TOOLS.md, or Copilot instructions.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:24
Finding

Untrusted Learning Promotion Can Poison Persistent Agent Instructions

Content
View full analysis
Remediation
View remediation

T01 · Skill Instruction Hijacking

Warning
Location
hooks/openclaw/handler.js:27
Finding

Opt-In Hooks Persistently Inject Skill Instructions into Agent Context

Content
View full analysis
{ // Safety checks for event structure if (!event || typeof event !== 'object') { return; } // Only handle agent:bootstrap events if (event.type !== 'agent' || event.action !== 'bootstrap') { return; } // Safety check for context if (!event.context || typeof event.context !== 'object') { return; } // Inject the reminder as a virtual bootstrap file // Check that bootstrapFiles is an array before pushing if (Array.isArray(event.context.bootstrapFiles)) { event.context.bootstrapFiles.push({ path: 'SELF_IMPROVEMENT_REMINDER.md', content: REMINDER_CONTENT, virtual: true, }); } }; ``` The injected content includes: ```javascript **Promote when pattern is proven:** - Behavioral patterns → \`SOUL.md\` - Workflow improvements → \`AGENTS.md\` - Tool gotchas → \`TOOLS.md\` ``` Installation and activation are documented at `SKILL.md:95-104`: ```bash # Copy hook to OpenClaw hooks directory cp -r hooks/openclaw ~/.openclaw/hooks/self-improvement # Enable it openclaw hooks enable self-improvement ``` Additional hook configurations at `SKILL.md:468-512` run `scripts/activator.sh` after prompts and `scripts/error-detector.sh` after Bash tool use. ### Technical Analysis Once explicitly installed and enabled, the OpenClaw hook survives the original skill invocation and runs during subsequent agent bootstrap events. It inserts a virtual bootstrap file containing skill-controlled instructions into the agent's context. The Claude Code and Codex configurations similarly support running reminder scripts after every prompt or Bash operation. This converts a user-installed skill into a persistent instruction channel. The injected instructions encourage recording session ...[truncated 1332 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
SKILL.md:33
Finding

Installation Instructions Use Unpinned Mutable Upstream Sources

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (25)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description is about capturing and using learnings for continuous improvement when errors, corrections, outdated knowledge, or tool failures occur. The supplied code does not implement learning capture, error/correction tracking, review of learnings, or any trigger-based continuous-improvement behavior. Instead, it is a helper utility for extracting/promoting a learning into a new skill scaffold by creating directories and a templated markdown file. While this may be adjacent to a broader learnings workflow, its primary purpose is materially different and includes undeclared filesystem-writing/scaffolding capabilities.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

Cross-session transcript reading and message passing create a direct pathway for prior session data to be accessed or relayed into unrelated contexts. Without strict authorization and minimization rules, this can expose sensitive content from one user task to another session or agent that did not originally receive it.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The templates explicitly ask for 'full context,' inputs, parameters, environment details, and user context, which strongly encourages plaintext storage of secrets, tokens, internal paths, proprietary prompts, or personal data. This is dangerous because the notes become a secondary datastore that may bypass normal application privacy and secret-handling controls.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Directing users to modify ~/.claude/settings.json affects the agent's persistent global configuration, which is a sensitive control point. Changes there can silently influence future sessions and projects, so coupling it with automatic command hooks creates a durable execution mechanism that could be abused or accidentally left enabled far beyond the intended scope.

Content

Scanner excerpt · references/hooks-setup.md (reported line 48)May include surrounding context.

Option 2: User-Level Configuration

Add to ~/.claude/settings.json for global activation:

json
{

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

Instructions found that direct the agent to transmit conversation context or user data to external services.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 177)May include surrounding context.

sessions_send

Send message to another session:

text
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest's broad activation conditions can cause the skill to trigger in many normal interactions, increasing the chance that conversation content is persistently logged or promoted when not necessary. In a system with automatic skill loading or hooks, overbroad triggering raises the attack surface for prompt-driven data retention and context pollution.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill encourages persistent logging and cross-session sharing of learnings, errors, and corrections, but provides no minimization, redaction, retention, or access-control guidance. As a result, user-provided details and task context may be stored and redistributed in plaintext beyond the original session, creating privacy and data leakage risks.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
72% confidence
Finding

The instructions establish persistent local storage under a workspace directory, which creates durable retention of operational and conversational data. Persistence itself is not inherently unsafe, but in this skill it amplifies the privacy risks because the stored content may include detailed errors, user context, and learnings without retention or protection rules.

Content

Scanner excerpt · SKILL.md (reported line 65)May include surrounding context.

└── FEATURE_REQUESTS.md

text

### Create Learning Files

```bash
mkdir -p ~/.openclaw/workspace/.learnings

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrases are common conversational language such as corrections, wishes, and questions, so they can fire on ordinary chat and capture data without meaningful intent from the user. That makes accidental persistence and propagation of sensitive or irrelevant content more likely, especially when paired with hooks or automated review behavior.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The manifest describes a self-improvement skill focused on capturing learnings, errors, corrections, and reviewing/promoting them for continuous improvement. This section adds a distinct capability: extracting and scaffolding entirely new skills via helper scripts and manual creation workflows, which is not a direct or necessary part of logging learnings.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The template asks authors to include trigger conditions, but it does not require those conditions to be concrete, bounded, or testable. In an agent skill system, vague activation criteria can cause over-broad invocation, leading the agent to apply the skill in unintended contexts and potentially run inappropriate guidance or helper commands.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The minimal template allows a description of 'what this skill does and when to use it' without requiring precise constraints, exclusions, or verification conditions. Minimal templates are likely to be copied as-is, so this omission can propagate ambiguous activation logic across many skills.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script-oriented template includes executable helpers but does not emphasize strict activation boundaries before those helpers are suggested or run. In practice, vague invocation criteria combined with scripts increases the risk that automation is applied in the wrong environment, against the wrong target, or without necessary preconditions.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · references/examples.md (reported line 301)May include surrounding context.

When the above learning is extracted as a skill, it becomes:

File: skills/docker-m1-fixes/SKILL.md

markdown
---

Session Persistence

Medium
Category
Rogue Agent
Confidence
78% confidence
Finding

Placing hook configuration in a project-local .claude/settings.json establishes persistent behavior for future sessions in that repository. While this is a legitimate feature, in this context it creates lasting automatic execution of local scripts, which can be risky if the repository is shared, the scripts change over time, or users do not realize the behavior persists.

Content

Scanner excerpt · references/hooks-setup.md (reported line 15)May include surrounding context.

Option 1: Project-Level Configuration

Create .claude/settings.json in your project root:

json
{

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

An empty matcher on a UserPromptSubmit hook causes the script to run on every prompt, creating a very broad automatic execution surface. In this skill context, that means any session interaction can trigger local scripts, amplifying the impact of mistakes, malicious modifications to the script, or unintended data exposure through hook processing.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The user-level configuration installs the hook in ~/.claude/settings.json for global activation, so the script will run across projects and contexts without meaningful constraints. That broad persistence increases the blast radius of any compromised or buggy script and can affect unrelated repositories, sessions, or sensitive workflows.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The Codex example also uses an empty matcher, so the hook's invocation conditions are effectively unconstrained. In a tool that can access project context and run with user permissions, unconditional hook execution unnecessarily expands exposure and makes accidental or malicious behavior more likely to trigger.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The document's security section asserts that the hook scripts only output text and do not run commands, but the same file configures them as command hooks and instructs users to execute an extraction script directly. That inconsistency can mislead users into underestimating the trust and execution risk of these scripts, increasing the chance they enable arbitrary local code execution with their agent's permissions.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
90% confidence
Finding

The instructions explicitly create a persistent .learnings/ store in the workspace or skill directory, enabling information from one interaction to survive into later sessions. In the context of a self-improvement skill, this persistence increases the blast radius of any accidental capture of sensitive data or any prompt-injected content that gets written as a 'learning.'

Content

Scanner excerpt · references/openclaw-integration.md (reported line 57)May include surrounding context.

openclaw hooks enable self-improvement

text

### 3. Create Learning Files

Create the `.learnings/` directory in your workspace:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The guide encourages storing learnings in persistent files and promoting them into workspace-wide prompt files, then sharing them across sessions, but it provides no warnings or controls for sensitive data. This creates a realistic risk that secrets, personal data, internal prompts, or incident details entered during failures and corrections will be retained and propagated beyond the original context.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The trigger definitions are broad enough that routine failures, vague user corrections, or normal model uncertainty could cause the skill to activate and write persistent records without clear user intent. In a self-improvement skill, this can lead to excessive or inaccurate logging, and can amplify prompt-injected or attacker-induced events into durable memory.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

Comments and usage indicate this tool 'creates a new skill from a learning entry,' giving the skill the capability to generate new project artifacts. For a skill whose stated purpose is capturing learnings, errors, and corrections, automatic creation of new skill directories and markdown manifests is a separate lifecycle-management capability not obviously required by that purpose.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest describes a skill focused on recording and reviewing learnings, corrections, and failures to support self-improvement. This script instead scaffolds entirely new skills by creating directories and writing a new SKILL.md template, which is a broader code-generation/project-modification behavior rather than simply capturing or reviewing learnings.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

Using placeholder trigger labels like 'Trigger 1' and 'Trigger 2' gives no model for specificity and normalizes generic invocation rules. This increases the chance that derived skills will be authored with ambiguous triggers, which can cause accidental or excessive skill activation.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.