Back to skill

Security audit

Self Improving Agent2

Security checks for vulnerabilities and agentic risk

Overview

This self-improvement skill is mostly transparent, but it can persist conversation-derived guidance into future agent instruction files and broad hooks without strong approval gates.

Install only if you want persistent agent self-improvement. Keep logs project-local where possible, do not enable global hooks unless you trust and review the scripts, avoid forwarding transcripts or raw command output, and require manual review before promoting any learning into CLAUDE.md, AGENTS.md, SOUL.md, TOOLS.md, or Copilot instructions.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T02 · Agent Memory Poisoning

Warning
Location
SKILL.md:357
Finding

Untrusted Conversation-Derived Rules Can Be Promoted into Persistent Agent Context

Content
View full analysis
= 3` - Seen across at least 2 distinct tasks - Occurred within a 30-day window Promotion targets: - `CLAUDE.md` - `AGENTS.md` - `.github/copilot-instructions.md` - `SOUL.md` / `TOOLS.md` for OpenClaw workspace-level guidance when applicable Write promoted rules as short prevention rules (what to do before/while coding), not long incident write-ups. ``` `SKILL.md:465`: ```markdown 7. **Promote aggressively** - if in doubt, add to CLAUDE.md or .github/copilot-instructions.md ``` `hooks/openclaw/handler.js:8-25`: ```javascript const REMINDER_CONTENT = ` ## Self-Improvement Reminder After completing tasks, evaluate if any learnings should be captured: **Log when:** - User corrects you → \`.learnings/LEARNINGS.md\` - Command/operation fails → \`.learnings/ERRORS.md\` - User wants missing capability → \`.learnings/FEATURE_REQUESTS.md\` - You discover your knowledge was wrong → \`.learnings/LEARNINGS.md\` - You find a better approach → \`.learnings/LEARNINGS.md\` **Promote when pattern is proven:** - Behavioral patterns → \`SOUL.md\` - Workflow improvements → \`AGENTS.md\` - Tool gotchas → \`TOOLS.md\` Keep entries simple: date, title, what happened, what to do differently. `.trim(); ``` `hooks/openclaw/handler.js:44-51`: ```javascript // Inject the reminder as a virtual bootstrap file // Check that bootstrapFiles is an array before pushing if (Array.isArray(event.context.bootstrapFiles)) { event.context.bootstrapFiles.push({ path: 'SELF_IMPROVEMENT_REMINDER.md', content: REMINDER_CONTENT, virtual: true, ...[truncated 3649 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (18)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description is about recording and reviewing learnings from failures, corrections, outdated knowledge, and better approaches. The supplied code does something materially different: it creates a new skill scaffold under a skills directory and writes a templated SKILL.md file. While the comments mention creating a skill from a learning entry, the script does not actually capture learnings, parse learning records, review prior learnings, or respond to the listed triggers such as command failures or user corrections. Its primary purpose is file/directory generation for skill extraction, which is an undeclared and materially different capability.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Referencing ~/.claude/settings.json directs users to place persistent executable hook configuration in a sensitive agent-wide config location. That is more dangerous than project-local setup because it affects all sessions and trusts files under a user-controlled directory that may be modified later, intentionally or accidentally.

Content

Scanner excerpt · references/hooks-setup.md (reported line 48)May include surrounding context.

Option 2: User-Level Configuration

Add to ~/.claude/settings.json for global activation:

json
{

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

The sessions_send example provides an explicit mechanism to transmit information to another session, which is an exfiltration-capable channel in multi-agent environments. Although the example shows benign content, the capability could be abused to move sensitive data, prompt injections, or hidden instructions across trust boundaries and persist them outside the original session.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 181)May include surrounding context.

sessions_send

Send message to another session:

text
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The description says to use the skill whenever a command fails unexpectedly, whenever the user corrects the agent, whenever a capability is missing, and also to review learnings before major tasks. These triggers are very broad and cover common conversational and development events without clear exclusion conditions, which increases the chance of routine interactions invoking the skill unnecessarily.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
92% confidence
Finding

The skill encourages persistent storage under ~/.openclaw/workspace/.learnings, which creates session-spanning retention of potentially sensitive operational details. Even though it warns not to log secrets and recommends sanitization, broad instructions to capture errors, corrections, and cross-session sharing increase the chance that sensitive context, paths, prompts, or redacted-but-still-sensitive metadata are retained longer than necessary.

Content

Scanner excerpt · SKILL.md (reported line 81)May include surrounding context.

└── FEATURE_REQUESTS.md

text

### Create Learning Files

```bash
mkdir -p ~/.openclaw/workspace/.learnings

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The listed triggers include broad phrases such as 'Actually, it should be...', 'Can you also...', 'Is there a way to...', and 'Why can't you...', which commonly occur in ordinary conversation. Because the section says to 'Automatically log when you notice' these patterns, it does not clearly limit logging to substantive or repeated learnings, making accidental activation more likely.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · references/examples.md (reported line 301)May include surrounding context.

When the above learning is extracted as a skill, it becomes:

File: skills/docker-m1-fixes/SKILL.md

markdown
---

Session Persistence

Medium
Category
Rogue Agent
Confidence
84% confidence
Finding

Placing hook configuration in project settings introduces session-persistent automatic behavior that survives across restarts and can affect anyone using the repository. In this skill's context, persistence makes the self-improvement mechanism more powerful but also more likely to create unreviewed prompt interception and repeated execution without ongoing user awareness.

Content

Scanner excerpt · references/hooks-setup.md (reported line 15)May include surrounding context.

Option 1: Project-Level Configuration

Create .claude/settings.json in your project root:

json
{

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

An empty matcher causes the hook to trigger on every prompt, creating a broad, always-on execution path. In a self-improvement skill, that increases the chance of unnecessary context injection, sensitive prompt handling, and persistent behavior the user may not fully expect on routine interactions.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The user-level example enables global automatic execution from the agent config directory with no trigger constraints, expanding the blast radius from one project to every future session. In combination with hooks, this creates persistent behavior across repositories and tasks, which can amplify mistakes, data exposure, or abuse if the referenced scripts are changed or replaced.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The Codex example also uses an empty matcher, so the hook executes on every prompt without contextual limitation. This broad triggering is risky because it normalizes blanket interception and reminder injection across all sessions rather than restricting behavior to relevant debugging or correction scenarios.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The document tells users to configure these shell scripts as hook commands, which means they are executed by the agent runtime, yet the security section states the scripts 'only output text' and 'don't modify files or run commands.' That is misleading because any invoked shell script inherently executes and can perform arbitrary side effects, causing users to underestimate the trust and review required before enabling them.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 57)May include surrounding context.

openclaw hooks enable self-improvement

text

### 3. Create Learning Files

Create the `.learnings/` directory in your workspace:

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The integration guide extends a narrowly scoped self-improvement skill into modifying higher-privilege persistent prompt files such as SOUL.md, TOOLS.md, and AGENTS.md. That creates a prompt-persistence channel where transient observations or poisoned inputs can be promoted into future behavioral, workflow, or tool instructions, increasing the chance of long-lived prompt injection or unsafe policy drift.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The documentation introduces cross-session transcript reading, messaging, and agent spawning capabilities that exceed the stated purpose of merely capturing learnings. Even with cautionary wording, exposing sessions_history and sessions_send in this skill context enables unnecessary access to potentially sensitive session data and creates channels for propagating tainted context across sessions.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The trigger conditions are broad enough to activate logging on common events such as knowledge gaps, API errors, or model surprises, which can lead to over-collection and persistence of low-quality or sensitive context. In a self-improvement system, this increases the chance that adversarial prompts, user secrets, or ephemeral mistakes are stored and later reinjected as trusted guidance.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The manifest describes a skill focused on recording learnings, errors, corrections, and reviewing those learnings before tasks. This script instead creates a new skill directory and generates a SKILL.md template, which is a separate skill-authoring/scaffolding operation rather than capturing or reviewing learnings themselves.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

For a skill described as capturing learnings, errors, and corrections, filesystem write operations that scaffold entirely new skills are not an obvious direct requirement. The code performs persistent project modification by creating directories and writing templated manifests, which goes beyond simple learning capture.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.