Back to skill

Security audit

Self Improving Agent

Security checks for vulnerabilities and agentic risk

Overview

This skill is not overtly destructive, but it asks agents to persist conversation-derived notes and promote them into future instruction files, with broad optional hooks.

Install only if you want persistent self-improvement logs and future-agent memory updates. Prefer project-local setup over global ~/.claude hooks, narrow any hook matcher, review hook scripts before enabling them, redact secrets and personal data before logging, and require human review before adding any learning to CLAUDE.md, AGENTS.md, SOUL.md, TOOLS.md, or copilot instruction files.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T02 · Agent Memory Poisoning

Warning
Location
SKILL.md:346
Finding

Untrusted Conversation Content Can Be Promoted into Persistent Agent Instructions

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 346–361
Vulnerability Type: Persistent agent memory and instruction poisoning
Risk Level: Medium

Vulnerable Code Snippet

markdown
### Promotion Rule (System Prompt Feedback)

Promote recurring patterns into agent context/system prompt files when all are true:

- `Recurrence-Count >= 3`
- Seen across at least 2 distinct tasks
- Occurred within a 30-day window

Promotion targets:
- `CLAUDE.md`
- `AGENTS.md`
- `.github/copilot-instructions.md`
- `SOUL.md` / `TOOLS.md` for OpenClaw workspace-level guidance when applicable

Write promoted rules as short prevention rules (what to do before/while coding),
not long incident write-ups.

Related guidance at SKILL.md, lines 448–448:

markdown
7. **Promote aggressively** - if in doubt, add to CLAUDE.md or .github/copilot-instructions.md

Technical Analysis

The skill directs agents to capture corrections and other conversation-derived material in .learnings/, then promote recurring entries into files that are loaded as persistent agent context. The documented promotion targets include CLAUDE.md, AGENTS.md, .github/copilot-instructions.md, SOUL.md, and TOOLS.md.

Conversation content is an untrusted input boundary. The workflow does not require human approval before promotion, distinguish trusted project facts from user-supplied directives, or filter entries containing role instructions, commands, URLs, tool-use directives, or attempts to weaken security constraints. Its recurrence threshold establishes frequency but not trustworthiness: an attacker can repeat a crafted correction across multiple tasks to satisfy the stated criteria.

Once promoted, the content can influence future sessions through normal workspace prompt loading. This constitutes an agent memory poisoning path. No automatic promotion implementation or embedded malicious payload was found, so exploitation depends on an agent following the documented ...[truncated 1773 chars]

Remediation
View remediation

Remediation Suggestions

  1. Require explicit human approval before writing any learning into persistent agent-context files.
  2. Treat conversation-derived entries as untrusted data, regardless of recurrence count or source claims.
  3. Replace automatic or aggressive promotion guidance with a review workflow that verifies:
    • The statement is a factual, project-specific rule.
    • The rule is supported by trusted repository documentation or code.
    • The entry does not contain role directives, safety overrides, commands, external URLs, credential instructions, or tool-routing instructions.
  4. Restrict direct promotion into high-impact files such as SOUL.md, TOOLS.md, and global workspace instructions. Prefer a human-reviewed candidate file that is not automatically loaded into agent context.
  5. Preserve provenance for every candidate, including source session, author, supporting files, reviewer, and approval timestamp.
  6. Use an allowlisted schema for promoted rules and render values as inert data rather than copying raw conversation text.
  7. Require code-owner or maintainer review for changes to persistent instruction files.
  8. Add duplicate-content and adversarial-pattern checks so repeated attacker input cannot become trusted merely by satisfying recurrence thresholds.
  9. Remove or revise the “Promote aggressively” instruction at line 448; promotion should be conservative and review-gated.
  10. Add regression tests demonstrating that corrections containing prompt-injection language cannot be promoted without explicit approval.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (20)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description is about capturing and reviewing learnings, errors, corrections, and failures for continuous improvement. The supplied code does not implement any learning capture, error logging, correction tracking, review workflow, or retrieval of prior learnings. Instead, it is a helper script for extracting/promoting a learning into a new skill scaffold by creating directories and writing a template file. This is a materially different primary purpose and includes undeclared filesystem write behavior. While the comments mention creating a skill from a learning entry, the code itself only scaffolds the skill structure and does not perform the declared learning-management behavior.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Writing to ~/.claude/settings.json establishes behavior in a sensitive agent configuration directory, affecting future sessions beyond the current project. In this context, the file is being used to install a persistent command hook, so compromise or misuse of that configuration can create durable, cross-project execution and data exposure risks.

Content

Scanner excerpt · references/hooks-setup.md (reported line 48)May include surrounding context.

Option 2: User-Level Configuration

Add to ~/.claude/settings.json for global activation:

json
{

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

Instructions found that direct the agent to transmit conversation context or user data to external services.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 177)May include surrounding context.

sessions_send

Send message to another session:

text
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The manifest description encourages invocation in a wide range of routine situations, including ordinary failures, corrections, and realization of outdated knowledge. In systems that auto-load or suggest skills from descriptions, this broad guidance can over-activate the skill and lead to unnecessary data capture and persistence.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 64)May include surrounding context.

└── FEATURE_REQUESTS.md

text

### Create Learning Files

```bash
mkdir -p ~/.openclaw/workspace/.learnings

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill explicitly encourages sharing learnings across sessions and reading other session transcripts, which creates a natural-language data exfiltration channel between contexts. If users discuss secrets, proprietary code, or personal data, those details can be propagated beyond the original session boundary without adequate need-to-know controls.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The logging templates instruct the agent to store full context, inputs, parameters, environment details, user requests, and related files. That creates a durable retention sink for sensitive information such as credentials, tokens, internal paths, customer data, or confidential prompts, especially when logging follows failures where raw error output often contains secrets.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The promotion workflow pushes learnings into long-term memory and instruction files intended to shape future agent behavior. Without strict sanitization, this can permanently embed sensitive user-derived information into persistent context, where it may later be surfaced, reused, or exposed across unrelated tasks.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrases are broad and map common conversational language to automatic logging actions. In agent environments with automatic activation or hooks, this can cause unintended persistence of ordinary user statements, corrections, or requests, increasing the chance that sensitive or irrelevant content is stored without deliberate review.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The best-practice guidance emphasizes logging immediately and promoting aggressively, but provides no compensating privacy controls. In practice, this increases the likelihood that sensitive data is retained reflexively before anyone evaluates whether it is necessary or safe to store.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · references/examples.md (reported line 301)May include surrounding context.

When the above learning is extracted as a skill, it becomes:

File: skills/docker-m1-fixes/SKILL.md

markdown
---

Session Persistence

Medium
Category
Rogue Agent
Confidence
84% confidence
Finding

Creating .claude/settings.json in the project root introduces session-persistent agent behavior that survives across prompts and restarts. Persistent hooks are not inherently unsafe, but in this skill they are used to execute shell commands automatically, so persistence materially increases the blast radius of mistakes, abuse, or later script tampering.

Content

Scanner excerpt · references/hooks-setup.md (reported line 15)May include surrounding context.

Option 1: Project-Level Configuration

Create .claude/settings.json in your project root:

json
{

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

An empty matcher causes the UserPromptSubmit hook to trigger on every prompt, creating an always-on execution path. In an agent environment, broad automatic triggering increases attack surface, raises the chance of prompt-driven abuse or unexpected side effects, and makes it harder for users to reason about when code will run.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This guidance combines two risky properties: user-level installation in ~/.claude/settings.json and an empty matcher that activates on all prompts globally. That creates persistent, broad-scope command execution across sessions and repositories, so any bug or malicious modification in the hooked script can affect all future agent use.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Although labeled 'minimal,' this setup still uses an empty matcher, so the hook executes for every prompt rather than only for relevant self-improvement events. That broad trigger scope can cause unnecessary code execution, information exposure through hook context, and unanticipated behavior in unrelated tasks.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The Codex CLI example repeats the empty matcher pattern, enabling the command hook for every prompt. In a coding-agent workflow, that means pervasive automatic script execution, which amplifies the consequences of script compromise, prompt abuse, or simple implementation mistakes.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The document claims the hook scripts 'only output text' and 'don't modify files or run commands,' but the configuration explicitly installs them as command hooks and even shows invoking another shell script. This misleading assurance can cause operators to grant trust or permissions they would not otherwise allow, increasing the chance of unsafe execution in a high-trust agent context.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
84% confidence
Finding

The integration explicitly creates a persistent .learnings/ directory, enabling session-derived information to survive resets and be reused later. In a self-improvement skill, that persistence materially increases the chance that sensitive user content, operational details, or incorrect guidance become durable state and influence future behavior beyond the original session.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 57)May include surrounding context.

openclaw hooks enable self-improvement

text

### 3. Create Learning Files

Create the `.learnings/` directory in your workspace:

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The 'Detection Triggers' section lists generic conditions such as 'Knowledge gaps', 'API errors', and user corrections like 'No, that's wrong...' without clearly constraining when the skill should activate or when it should not. These phrases overlap with common interaction patterns and lack negative examples or scope boundaries, making unintended invocation more likely.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The setup instructions direct users to create persistent learning storage in workspace or skill directories without clearly warning that future runs may write durable data there. This can surprise users, lead to retention of sensitive prompts or errors, and increase privacy risk in environments where workspace contents are synced or shared.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.