Back to skill

Security audit

Self Improving Agent 3.0.2

Security checks for vulnerabilities and agentic risk

Overview

This self-improvement skill is mostly transparent about what it does, but it can persist conversation-derived learnings into future agent instructions and install recurring hooks with broad scope.

Install only if you want persistent self-improvement memory. Keep hooks project-scoped when possible, review hook scripts before enabling them, avoid global activation unless you trust the source, and require human review/redaction before any learning is written to shared or auto-loaded instruction files.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T01 · Skill Instruction Hijacking

Error
Location
hooks/openclaw/handler.js:9
Finding

Automatic Injection of Skill-Controlled Instructions into Agent Context

Content
View full analysis
After completing this task, evaluate if extractable knowledge emerged: - Non-obvious solution discovered through investigation? - Workaround for unexpected behavior? - Project-specific pattern learned? - Error required debugging to resolve? If yes: Log to .learnings/ using the self-improvement skill format. If high-value (recurring, broadly applicable): Consider skill extraction. EOF ``` ### Technical Analysis The OpenClaw hook adds skill-controlled content to `bootstrapFiles`, which are supplied to the agent during bootstrap. The shell hook similarly emits instructional content whenever the configured `UserPromptSubmit` event occurs. Consequent ...[truncated 1813 chars]
Remediation
View remediation

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:347
Finding

Untrusted Learnings Can Be Promoted into Persistent Agent Instruction Files

Content
View full analysis
= 3` - Seen across at least 2 distinct tasks - Occurred within a 30-day window Promotion targets: - `CLAUDE.md` - `AGENTS.md` - `.github/copilot-instructions.md` - `SOUL.md` / `TOOLS.md` for OpenClaw workspace-level guidance when applicable Write promoted rules as short prevention rules (what to do before/while coding), not long incident write-ups. ``` The broader promotion workflow identifies persistent targets: ```markdown | Behavioral patterns | Promote to `SOUL.md` | | Workflow improvements | Promote to `AGENTS.md` | | Tool gotchas | Promote to `TOOLS.md` | ``` ### Technical Analysis The skill captures information originating from user corrections, conversations, command failures, API failures, and agent observations. It then directs the agent to promote recurring entries into files that OpenClaw or other coding agents load as persistent instructions. Recurrence is not a sufficient trust or correctness control. An attacker can repeat the same malicious or misleading claim across tasks, or cause crafted command output to be recorded repeatedly. Once distilled into `CLAUDE.md`, `AGENTS.md`, `.github/copilot-instructions.md`, `SOUL.md`, or `TOOLS.md`, that claim can affect subsequent sessions as an instruction rather than remaining isolated as untrusted historical data. The project does not define mandatory human approval, provenance verification, instruction-content filtering, or a policy that prevents raw attacker-influenced material from becoming executable agent guidance. ...[truncated 1444 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
SKILL.md:39
Finding

Manual Installation Uses an Unpinned Mutable Git Repository

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (19)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description is about maintaining and consulting a repository of learnings when errors, corrections, outdated knowledge, or better approaches are discovered. The supplied code does not implement capture, storage, retrieval, or review of learnings. Instead, it is a utility that generates a new skill folder and a templated SKILL.md file based on a skill name, optionally as a dry run. Its primary purpose is skill scaffolding/promoting a learning into a skill artifact, which is materially different from the declared behavior. The file-creation capability is also undeclared in the description and permissions.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Directing users to place hook configuration in ~/.claude/settings.json establishes persistent execution from an agent config directory, which is a high-value location because settings there affect all sessions. If an attacker can influence the referenced script path or convince a user to install an unsafe skill, the resulting execution becomes durable and broadly impactful.

Content

Scanner excerpt · references/hooks-setup.md (reported line 48)May include surrounding context.

Option 2: User-Level Configuration

Add to ~/.claude/settings.json for global activation:

json
{

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

Instructions found that direct the agent to transmit conversation context or user data to external services.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 177)May include surrounding context.

sessions_send

Send message to another session:

text
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The description says to use the skill when a command fails unexpectedly, when the user corrects Claude, when a better approach is discovered, and to review learnings before major tasks. These conditions are broad and likely to overlap with many ordinary interactions, without clear exclusion criteria or negative examples to limit when the skill should or should not activate.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly instructs the agent to persist user corrections, requests, and error context into local or workspace markdown files, which can capture sensitive conversational content, credentials, proprietary prompts, or personal data. Because these logs are meant for future reuse, the risk is not just temporary exposure but durable retention and later disclosure to other sessions or collaborators.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

Promoting conversation-derived learnings into persistent instruction files like CLAUDE.md, AGENTS.md, or Copilot instructions spreads information beyond the original session boundary and can influence future agent behavior. If sensitive details or user-specific context are promoted, they become a long-lived disclosure source and may be surfaced to unrelated tasks, users, or repositories.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
87% confidence
Finding

The skill recommends creating persistent storage under ~/.openclaw/workspace/.learnings, which establishes session-to-session retention of potentially sensitive operational and conversational data. Persistence itself is not malicious, but in this skill’s context it increases exposure duration and broadens who or what may later access the stored information.

Content

Scanner excerpt · SKILL.md (reported line 64)May include surrounding context.

└── FEATURE_REQUESTS.md

text

### Create Learning Files

```bash
mkdir -p ~/.openclaw/workspace/.learnings

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The documented cross-session tools enable reading transcripts and sending learnings between sessions, which materially increases the chance of natural-language data leakage across otherwise separate contexts. In a multi-agent or shared workspace environment, this can expose confidential prompts, incident details, credentials in logs, or project-sensitive reasoning to other sessions without a clear need-to-know boundary.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The templates ask for full context, parameters, error output, and user context, creating a strong incentive to store secrets and sensitive data verbatim. Since command failures and debugging often include environment variables, API responses, file paths, or stack traces containing confidential information, these templates can become a high-value leakage sink.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The 'Detection Triggers' section includes phrases like 'Can you also...', 'Is there a way to...', and broad conditions such as 'Unexpected output or behavior.' These are common in everyday chat and could cause unintended invocation or over-logging because the file does not sufficiently constrain context or provide negative examples.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · references/examples.md (reported line 301)May include surrounding context.

When the above learning is extracted as a skill, it becomes:

File: skills/docker-m1-fixes/SKILL.md

markdown
---

Session Persistence

Medium
Category
Rogue Agent
Confidence
83% confidence
Finding

Project-level settings create session-persistent automatic behavior that survives beyond a single interaction, so anyone using the repository may unknowingly inherit the hook execution. In this context, the persistence is more significant because the skill is designed to trigger proactively and can affect every future session in that project.

Content

Scanner excerpt · references/hooks-setup.md (reported line 15)May include surrounding context.

Option 1: Project-Level Configuration

Create .claude/settings.json in your project root:

json
{

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

An empty matcher causes the hook to fire on every prompt, creating broad automatic execution with no scope restriction. In this skill context, that means a local script is invoked for all user interactions, increasing exposure to prompt-triggered behavior, performance overhead, and abuse if the script or repo is modified.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The user-level configuration installs the hook in ~/.claude/settings.json for global activation across sessions, which broadens the trust boundary from one project to all uses of the agent. In a self-improvement skill that injects reminders automatically, this persistence makes accidental or unintended execution more dangerous if the referenced script changes or comes from an untrusted location.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The Codex example repeats the empty matcher pattern, again causing execution on every prompt without contextual restriction. Because the file is a setup guide, users may copy-paste this configuration directly, propagating overly broad hook behavior into another agent environment.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The document's security section is internally inconsistent: the configured hooks explicitly execute shell scripts as commands, yet it claims the scripts 'only output text' and 'don't modify files or run commands.' This can mislead users into underestimating the execution and privilege risks of enabling the hooks, especially because the scripts run with the agent's permissions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The file describes workspace-based prompt injection as a feature without warning that workspace files are untrusted input and can manipulate model behavior. This is dangerous because users may treat injected files like AGENTS.md, SOUL.md, or TOOLS.md as authoritative, allowing malicious or tampered workspace content to steer the agent into unsafe actions or data disclosure.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
80% confidence
Finding

The guide instructs persistent storage of learnings in workspace or skill directories without discussing retention, secret handling, or file permission controls. Persistent memory can accumulate credentials, proprietary prompts, errors containing sensitive payloads, or personal data, making later compromise or inadvertent reuse more likely.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 57)May include surrounding context.

openclaw hooks enable self-improvement

text

### 3. Create Learning Files

Create the `.learnings/` directory in your workspace:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The document encourages cross-session transcript access and messaging but provides no guidance on handling sensitive data, consent, minimization, or access boundaries. In an agent environment, session history can contain secrets, user data, or privileged context, so normalizing unrestricted access increases the risk of privacy leakage and unintended data sharing.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.