Back to skill

Security audit

Self Improving Agent

Security checks for vulnerabilities and agentic risk

Overview

The skill is not clearly malicious, but it encourages broad persistent agent-memory changes and always-on hooks without enough user control or redaction guidance.

Install only if you are comfortable with persistent learning files and hook-based reminders. Keep .learnings private by default, redact secrets and personal or customer data, avoid global always-on hooks, and require explicit review before promoting any learning into agent instruction files such as AGENTS.md, SOUL.md, TOOLS.md, CLAUDE.md, or Copilot instructions.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:16
Finding

Untrusted conversation content can poison persistent agent instructions

Content
View full analysis
= 3` ``` `SKILL.md:448`: ```markdown 7. **Promote aggressively** - if in doubt, add to CLAUDE.md or .github/copilot-instructions.md ``` `hooks/openclaw/handler.js:14-23`: ```javascript - User corrects you → \`.learnings/LEARNINGS.md\` - Command/operation fails → \`.learnings/ERRORS.md\` - User wants missing capability → \`.learnings/FEATURE_REQUESTS.md\` - You discover your knowledge was wrong → \`.learnings/LEARNINGS.md\` - You find a better approach → \`.learnings/LEARNINGS.md\` **Promote when pattern ...[truncated 2490 chars]
Remediation
View remediation

T01 · Skill Instruction Hijacking

Warning
Location
scripts/activator.sh:8
Finding

Opt-in hooks persistently inject Skill-controlled instructions into agent context

Content
View full analysis
After completing this task, evaluate if extractable knowledge emerged: - Non-obvious solution discovered through investigation? - Workaround for unexpected behavior? - Project-specific pattern learned? - Error required debugging to resolve? If yes: Log to .learnings/ using the self-improvement skill format. If high-value (recurring, broadly applicable): Consider skill extraction. EOF ``` `hooks/openclaw/handler.js:30-52`: ```javascript const handler = async (event) => { // Safety checks for event structure if (!event || typeof event !== 'object') { return; } // Only handle agent:bootstrap events if (event.type !== 'agent' || event.action !== 'bootstrap') { return; } // Safety check for context if (!event.context || typeof event.context !== 'object') { return; } // Inject the reminder as a virtual bootstrap file // Check that bootstrapFiles is an array before pushing if (Array.isArray(event.context.bootstrapFiles)) { event.context.bootstrapFiles.push({ path: 'SELF_IMPROVEMENT_REMINDER.md', content: REMINDER_CONTENT, virtual: true, }); } }; ``` `SKILL.md:477-510`: ```json { "hooks": { "UserPromptSubmit": [{ "matcher": "", "hooks": [{ "type": "command", "command": "./skills/self-improvement/scripts/activator.sh" }] }] } } ``` ```json { "hooks": { "UserPromptSubmit": [{ "matcher": "", "hooks": [{ "type": "command", "command": "./skills/self-improvement/scripts/activato ...[truncated 2443 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:174
Finding

Error-learning workflow can retain and expose sensitive command output

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (22)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description is about recording and using learnings from failures, corrections, outdated knowledge, and better approaches. The supplied code does not capture errors, corrections, or learnings, nor does it review prior learnings before tasks. Instead, it is a filesystem-oriented utility that creates a new skill scaffold (directory plus SKILL.md template) based on a skill name, optionally in dry-run mode. While it references 'learning entry' and 'promoted_to_skill', that is only contextual metadata for template generation. The primary purpose and concrete behavior are materially different from the declared purpose, so this is a clear mismatch.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
93% confidence
Finding

Guidance to modify ~/.claude/settings.json instructs users to alter a persistent agent configuration directory in their home folder. Because that config can automatically execute hook commands in future sessions, this creates a durable trust anchor and raises the impact of misconfiguration, compromise, or unnoticed behavior changes.

Content

Scanner excerpt · references/hooks-setup.md (reported line 48)May include surrounding context.

Option 2: User-Level Configuration

Add to ~/.claude/settings.json for global activation:

json
{

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

Instructions found that direct the agent to transmit conversation context or user data to external services.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 177)May include surrounding context.

sessions_send

Send message to another session:

text
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill persistently logs learnings, errors, and user corrections to markdown files but does not warn that those logs may contain sensitive information. In practice, operators may use it in contexts involving secrets, internal paths, stack traces, or personal data, causing unintentional retention of sensitive content on disk.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
90% confidence
Finding

The skill instructs creation of persistent learning files in a workspace directory, establishing durable memory across sessions. Persistent storage is not inherently malicious, but in this skill it materially increases risk because other sections encourage writing errors, inputs, and user context into those files.

Content

Scanner excerpt · SKILL.md (reported line 64)May include surrounding context.

└── FEATURE_REQUESTS.md

text

### Create Learning Files

```bash
mkdir -p ~/.openclaw/workspace/.learnings

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The inter-session communication section encourages reading other sessions' transcripts and sending learnings across sessions, which can spread sensitive user content beyond the original context. That expands data exposure and retention boundaries without any stated need-to-know checks, minimization rules, or redaction requirements.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The learning-entry format asks for full context and related details in persistent markdown, which can easily capture secrets, proprietary code context, customer data, or internal operational information. Because these entries are intended for later review and possible promotion, the data may be retained and redistributed well beyond the original task.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
99% confidence
Finding

The error-entry instructions explicitly tell the agent to store actual error messages, command context, and input parameters, all of which commonly contain secrets, access tokens, file paths, hostnames, and internal configuration. Persisting that material in plain-text markdown creates a direct confidentiality risk and a durable source of accidental disclosure.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The listed trigger phrases such as "Can you also...", "I wish you could...", and "Is there a way to..." are very common conversational patterns and are not specific to this skill's intended logging behavior. Without tighter scope or exclusion examples, they could cause the skill to activate in many ordinary interactions unrelated to self-improvement logging.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

The instruction to apply the skill whenever the agent discovers something non-obvious, corrects itself, learns conventions, hits errors, or finds better approaches is expansive and leaves unclear when the skill should not be invoked. This ambiguity increases the chance of unintended or excessive activation during routine work.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · references/examples.md (reported line 301)May include surrounding context.

When the above learning is extracted as a skill, it becomes:

File: skills/docker-m1-fixes/SKILL.md

markdown
---

Session Persistence

Medium
Category
Rogue Agent
Confidence
80% confidence
Finding

The guide instructs users to create persistent hook configuration in .claude/settings.json, causing the behavior to survive across sessions. Session persistence is not inherently malicious, but in this context it matters because it keeps executable hooks active automatically and may outlast user awareness.

Content

Scanner excerpt · references/hooks-setup.md (reported line 15)May include surrounding context.

Option 1: Project-Level Configuration

Create .claude/settings.json in your project root:

json
{

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

Using an empty matcher causes the hook to fire on every prompt, maximizing exposure and making the self-improvement script part of all interactions. In a skill that installs executable hooks, broad automatic triggering increases the chance of sensitive-context capture, prompt pollution, or abuse if the script behavior changes or is compromised.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The user-level configuration recommends global activation in the user's home config, which broadens scope across all projects and sessions without meaningful constraints. That persistence amplifies the effect of any unsafe script behavior or future modification because the hook will run automatically in many contexts.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The Codex example repeats the empty matcher pattern, creating all-prompt activation with no limiting condition. In agent environments, always-on hook execution is risky because it expands the attack surface and can unintentionally process sensitive or irrelevant prompts.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The security section makes a misleading safety claim: these scripts are in fact executed as hook commands, and the guide also shows direct shell execution of one script. That can cause users to underestimate the trust boundary and install hooks assuming they are inert text emitters, when they actually run with the agent's permissions.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
72% confidence
Finding

Creating a persistent .learnings/ directory establishes retention of model-generated content across sessions. Persistence is not automatically unsafe, but in this skill context it increases the blast radius of poisoned or sensitive content because stored learnings can later influence behavior or reveal prior interactions.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 57)May include surrounding context.

openclaw hooks enable self-improvement

text

### 3. Create Learning Files

Create the `.learnings/` directory in your workspace:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The guide instructs the system to log learnings into workspace files and promote them into durable prompt documents without warning that this modifies user-controlled workspace content. Silent or poorly disclosed writes to workspace state can create persistence, overwrite user intent, and establish a prompt-injection foothold for future sessions.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The guide expands a narrowly scoped self-improvement skill into editing persistent high-privilege prompt files such as AGENTS.md, SOUL.md, and TOOLS.md. That turns transient 'learning capture' into durable behavioral reconfiguration, creating a prompt-persistence channel that can silently alter future agent behavior across sessions.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The documentation introduces cross-session transcript access, messaging, and task spawning even though the skill's stated purpose is local learning capture. Those capabilities broaden the trust boundary and can expose unrelated session data or propagate untrusted content between sessions, which is unnecessary for the described function.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The trigger conditions are broad and subjective, such as 'knowledge gaps' and 'model behavior surprise,' which can cause the skill to activate in many ambiguous situations. Over-inclusive activation increases the chance of recording or promoting sensitive, incorrect, or attacker-influenced content into persistent memory or prompt files.

Content

No source excerpt is available for this finding.

Ssd 3

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

Feature-request logging asks for user context and problem details without sensitivity boundaries, which can lead to storing personal, business, or security-relevant information that was only needed transiently during the conversation. Even if less risky than raw error logs, this still creates unnecessary retention of user-provided data.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.