Back to skill

Security audit

self_improving_agent

Security checks for vulnerabilities and agentic risk

Overview

The skill is a disclosed learning logger, but it encourages always-on hooks and durable changes to agent instruction files that can influence future sessions.

Install only if you want a persistent learning system that can influence future agent sessions. Keep hooks project-scoped, avoid empty matchers for routine use, review every promotion into CLAUDE.md, AGENTS.md, SOUL.md, TOOLS.md, or copilot instructions, and do not store secrets, raw transcripts, credentials, or sensitive command output in learning files.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (2)

T01 · Skill Instruction Hijacking

Error
Location
scripts/activator.sh:7
Finding

Automatic Bootstrap and Prompt Hook Instruction Injection

Content
View full analysis
After completing this task, evaluate if extractable knowledge emerged: - Non-obvious solution discovered through investigation? - Workaround for unexpected behavior? - Project-specific pattern learned? - Error required debugging to resolve? If yes: Log to .learnings/ using the self-improvement skill format. If high-value (recurring, broadly applicable): Consider skill extraction. EOF ``` `hooks/openclaw/handler.js:39-46`: ```javascript // Inject the reminder as a virtual bootstrap file // Check that bootstrapFiles is an array before pushing if (Array.isArray(event.context.bootstrapFiles)) { event.context.bootstrapFiles.push({ path: 'SELF_IMPROVEMENT_REMINDER.md', content: REMINDER_CONTENT, virtual: true, }); } ``` `hooks/openclaw/handler.ts:54-61`: ```typescript // Inject the reminder as a virtual bootstrap file // Check that bootstrapFiles is an array before pushing if (Array.isArray(event.context.bootstrapFiles)) { event.context.bootstrapFiles.push({ path: 'SELF_IMPROVEMENT_REMINDER.md', content: REMINDER_CONTENT, virtual: true, }); } ``` ### Technical Analysis The package supplies hooks that inject skill-controlled instructions into agent context. The shell activator prints an instruction block intended to be consumed as system context after each matching prompt. The OpenClaw handler adds a virtual Markdown file to `bootstrapFiles`, causing its content to be loaded during agent bootstrap. The behavior is opt-in and documented, but once the hook is enabled, the injected instructions apply beyond an e ...[truncated 1686 chars]
Remediation
View remediation

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:330
Finding

Persistent Agent Instruction Poisoning Through Learning Promotion

Content
View full analysis
" .learnings/LEARNINGS.md` 4. If found: - Increment `Recurrence-Count` - Update `Last-Seen` - Add `See Also` links to related entries/tasks 5. If not found: - Create a new `LRN-...` entry - Set `Source: simplify-and-harden` - Set `Pattern-Key`, `Recurrence-Count: 1`, and `First-Seen`/`Last-Seen` ### Promotion Rule (System Prompt Feedback) Promote recurring patterns into agent context/system prompt files when all are true: - `Recurrence-Count >= 3` - Seen across at least 2 distinct tasks - Occurred within a 30-day window Promotion targets: - `CLAUDE.md` - `AGENTS.md` - `.github/copilot-instructions.md` - `SOUL.md` / `TOOLS.md` for OpenClaw workspace-level guidance when applicable Write promoted rules as short prevention rules (what to do before/while coding), not long incident write-ups. ``` `SKILL.md:440-448`: ```markdown ## Best Practices 1. **Log immediately** - context is freshest right after the issue 2. **Be specific** - future agents need to understand quickly 3. **Include reproduction steps** - especially for errors 4. **Link related files** - makes fixes easier 5. **Suggest concrete fixes** - not just "investigate" 6. **Use consistent categories** - enables filtering 7. **Promote aggressively** - if in doubt, add to CLAUDE.md or .github/copilot-instructions.md 8. **Review regularly** - stale ...[truncated 3259 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (23)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description is about a reflective learning/continuous-improvement skill that records errors, corrections, and lessons learned. The supplied code does not capture, review, or manage learnings directly. Instead, it is a helper utility for extracting/promoting a learning into a new skill by creating a local folder and templated SKILL.md file. That is a materially different primary purpose and includes undeclared filesystem-writing capabilities. While related at a high level to learnings, the actual behavior is specifically skill scaffolding, not learning capture or review.

Content

No source excerpt is available for this finding.

Vague Triggers

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The hook configuration uses empty matchers, effectively causing activation on every prompt and relevant tool event. This creates an unbounded automatic execution path for reminder and error-detection scripts, magnifying the risk of prompt-induced persistence, noisy logging, and unintended side effects across all sessions.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Directing users to modify ~/.claude/settings.json establishes persistent behavior in the agent's config directory, which is a sensitive control point for future sessions. In the context of executable command hooks, this becomes more dangerous because it grants ongoing code execution/influence across projects rather than a one-off local change.

Content

Scanner excerpt · references/hooks-setup.md (reported line 48)May include surrounding context.

Option 2: User-Level Configuration

Add to ~/.claude/settings.json for global activation:

json
{

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

Instructions found that direct the agent to transmit conversation context or user data to external services.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 177)May include surrounding context.

sessions_send

Send message to another session:

text
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The activation description is extremely broad and can cause the skill to engage during many ordinary interactions, increasing the chance that it logs unnecessary conversation content or modifies persistent memory files too often. In agent systems, overbroad activation expands the attack surface for prompt-triggered persistence and context pollution.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 64)May include surrounding context.

└── FEATURE_REQUESTS.md

text

### Create Learning Files

```bash
mkdir -p ~/.openclaw/workspace/.learnings

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This section encourages sharing learnings across sessions and interacting with other session transcripts, which can propagate user-provided corrections, context, and possibly sensitive material into broader long-term storage. Cross-session persistence substantially increases confidentiality risk because information may outlive the original context and be exposed to future tasks or agents.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The logging formats direct the agent to retain full error messages, input/parameters, user context, and detailed incident descriptions in persistent files. Those fields commonly contain secrets, personal data, tokens, internal paths, or proprietary business context, so the workflow creates a direct natural-language data leakage channel.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrases are common conversational language like corrections and wishes, so routine chat can unintentionally activate persistence workflows. That makes prompt injection easier because an attacker can embed these phrases to coerce the agent into recording or propagating attacker-controlled content.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The instruction to 'promote aggressively' encourages moving conversation-derived information into broad, persistent context files that influence future sessions. This increases both confidentiality risk and prompt-poisoning risk, because a bad or sensitive learning can become durable global guidance.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The agent-agnostic guidance relies on subjective conditions like 'discover something non-obvious' or 'find better approaches,' which are too ambiguous for safe automated activation. Ambiguity encourages overcollection and inconsistent persistence behavior, especially in agents that consume instructions literally.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · references/examples.md (reported line 301)May include surrounding context.

When the above learning is extracted as a skill, it becomes:

File: skills/docker-m1-fixes/SKILL.md

markdown
---

Session Persistence

Medium
Category
Rogue Agent
Confidence
80% confidence
Finding

Creating .claude/settings.json in the project root introduces persistent session behavior for anyone using that repository, which can silently alter future agent runs. Because the file configures command hooks, persistence matters: users may inherit automatic script execution without fully realizing it, especially in shared or cloned projects.

Content

Scanner excerpt · references/hooks-setup.md (reported line 15)May include surrounding context.

Option 1: Project-Level Configuration

Create .claude/settings.json in your project root:

json
{

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

An empty matcher causes the hook to run on every prompt, greatly expanding the number of opportunities for the script to influence agent behavior or process sensitive context. In a self-improvement skill, broad triggering is especially risky because it normalizes constant injection into the session and increases exposure of prompt contents to hook logic.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The user-level configuration installs a globally active hook in ~/.claude/settings.json with an unconstrained matcher, creating persistent execution across all projects and sessions. This broad scope magnifies the blast radius of any bug, script tampering, or unexpected data handling because every prompt can trigger the hook under the user's account.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The Codex CLI example repeats the same empty matcher pattern, propagating an overly broad hook configuration to another toolchain. This encourages always-on execution and increases the chance that sensitive prompts or unrelated workflows are processed by the hook scripts.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The guide states that the scripts 'only output text' and 'don't modify files or run commands', but the documented mechanism is a command hook that executes shell scripts. That mismatch can cause users to underestimate the trust boundary and install code execution hooks without appropriate scrutiny, increasing the risk of arbitrary code execution if the scripts are changed or replaced.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 57)May include surrounding context.

openclaw hooks enable self-improvement

text

### 3. Create Learning Files

Create the `.learnings/` directory in your workspace:

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The promotion workflow instructs writing persistent behavioral, workflow, and tool guidance into AGENTS.md, SOUL.md, and TOOLS.md, which are injected into future sessions. That creates a durable prompt-modification channel beyond simple error logging, allowing accidental or adversarial content to influence later agent behavior across the workspace.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The guide explicitly documents cross-session transcript access, session spawning, and inter-session messaging, which extends the skill from local learning capture into broader coordination and data-sharing capabilities. In a prompt-injection-prone workspace model, this increases the chance that sensitive context or learned content is propagated across sessions without clear scoping or authorization boundaries.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger conditions are broad terms like 'knowledge gaps,' 'model behavior surprise,' and 'skill issue,' which can fire in many normal interactions. Ambiguous activation criteria can cause over-collection, unnecessary persistence, and unintended modification of shared prompt files, especially in systems already using workspace prompt injection.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

Even the 'minimal setup' still uses an empty matcher, so the hook remains active for every prompt despite presenting itself as lower overhead. While the impact is lower than the global configuration, it still broadens prompt interception unnecessarily and may collect or influence unrelated interactions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The SOUL.md example instructs the agent to 'Avoid unnecessary caveats and disclaimers' as a blanket communication rule. This is a natural-language behavioral directive that may suppress safety or policy-relevant disclosures without offering a user choice or a documented, context-specific justification.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.