Back to skill

Security audit

Self Improving Agent 3.0.6

Security checks for vulnerabilities and agentic risk

Overview

This skill has a coherent learning-log purpose, but it also pushes conversation-derived notes into future agent instructions and optional automatic hooks without enough safeguards.

Review this carefully before installing. Keep learning files local by default, avoid recording secrets or private transcript content, require explicit human approval before promoting anything into agent instruction or memory files, and do not enable the hook setup unless the exact scripts are bundled or pinned and independently reviewed.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:19
Finding

Untrusted Conversation Content Can Poison Persistent Agent Instructions

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:19-27, 75-82, 145-158, 263-292, 348-358, 448
Vulnerability Type: Persistent agent memory poisoning
Risk Level: High

Vulnerable Code Snippet

markdown
| Situation | Action |
|-----------|--------|
| User corrects you | Log to `.learnings/LEARNINGS.md` with category `correction` |
...
| Broadly applicable learning | Promote to `CLAUDE.md`, `AGENTS.md`, and/or `.github/copilot-instructions.md` |
| Behavioral patterns | Promote to `SOUL.md` (OpenClaw workspace) |
markdown
### Metadata
- Source: conversation | error | user_feedback
- Related Files: path/to/file.ext
- Tags: tag1, tag2
markdown
Promote recurring patterns into agent context/system prompt files when all are true:

- `Recurrence-Count >= 3`
- Seen across at least 2 distinct tasks
- Occurred within a 30-day window

Promotion targets:
- `CLAUDE.md`
- `AGENTS.md`
- `.github/copilot-instructions.md`
- `SOUL.md` / `TOOLS.md` for OpenClaw workspace-level guidance when applicable
markdown
7. **Promote aggressively** - if in doubt, add to CLAUDE.md or .github/copilot-instructions.md

Technical Analysis

The skill directs the agent to derive learning records from conversation content and user feedback, then promote those records into persistent files such as CLAUDE.md, AGENTS.md, .github/copilot-instructions.md, and SOUL.md. These files can be automatically loaded as agent context in future sessions.

User-provided corrections and conversation text are not inherently trusted. The documented workflow does not require an independent verification step, explicit human approval, provenance-based trust decision, or security review before converting such content into durable instructions. Although one section describes a recurrence threshold, the separate instruction to “promote aggressively” and promote “if in doubt” weakens that safeguard.

...[truncated 1516 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove the “promote aggressively” and “if in doubt” guidance. Default to not promoting uncertain content.
  2. Require explicit human approval before modifying any automatically loaded instruction or memory file.
  3. Never promote conversation content verbatim. Convert verified facts into narrowly scoped, declarative rules.
  4. Require corroboration from trusted project documentation, reviewed code, tests, or an approved maintainer.
  5. Prohibit promotion of instructions that alter security controls, permissions, tool policies, credential handling, network access, or review requirements.
  6. Record provenance for every promoted rule, including the originating session, approving reviewer, supporting evidence, and review date.
  7. Treat user feedback and external text as untrusted input and scan it for embedded instructions before storage.
  8. Provide a review and rollback mechanism for all persistent agent-context changes.

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
SKILL.md:88
Finding

Cross-Session Transcript Access and Diagnostic Logging Lack Data-Minimization Controls

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:88-93, 182-191, 455-459
Vulnerability Type: Excessive access to session and environment data
Risk Level: Medium

Vulnerable Code Snippet

markdown
### Inter-Session Communication

OpenClaw provides tools to share learnings across sessions:

- **sessions_list** — View active/recent sessions
- **sessions_history** — Read another session's transcript  
- **sessions_send** — Send a learning to another session
- **sessions_spawn** — Spawn a sub-agent for background work
markdown
### Context
- Command/operation attempted
- Input or parameters used
- Environment details if relevant
markdown
**Track learnings in repo** (team-wide):
Don't add to .gitignore - learnings become shared knowledge.

Technical Analysis

The skill recommends reading other session transcripts, transmitting learnings between sessions, and preserving command inputs and environment details. It also presents repository tracking as an option for sharing those records.

The workflow does not define a need-to-know boundary, consent requirement, secret-redaction process, retention period, or restrictions on copying private transcript content. Command parameters, environment details, and transcripts can contain access tokens, credentials, internal paths, private user data, proprietary source material, or sensitive operational information.

Reading complete transcripts is broader than necessary for local error logging. Persisting extracted information in repository-tracked Markdown files or sending it to another session increases the number of principals and contexts that can access it.

Attack Path

  1. A user or tool invocation includes a secret, private value, internal path, or sensitive command parameter.
  2. An operation fails or produces an unexpected result.
  3. Following the skill, the agent reads the originating or another session’s transcript a ...[truncated 939 chars]
Remediation
View remediation

Remediation Suggestions

  1. Remove cross-session transcript access from the default workflow.
  2. Require explicit user consent and a documented need before reading another session’s history.
  3. Retrieve only the minimum relevant excerpt rather than an entire transcript.
  4. Explicitly prohibit recording passwords, API keys, tokens, cookies, private keys, personal data, and raw environment-variable values.
  5. Add automatic redaction for credentials, authorization headers, connection strings, sensitive command parameters, and common secret formats.
  6. Require confirmation before sending a learning to another session or committing it to a repository.
  7. Define retention and deletion policies for local and shared learning files.
  8. Store sensitive diagnostics in access-controlled local storage rather than repository-tracked Markdown files.
  9. Add a pre-commit check that detects likely secrets in .learnings/.

T08 · Insecure Dependencies

Warning
Location
SKILL.md:38
Finding

Unpinned Remote Repository Can Supply Persistently Executed Hook Scripts

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:38-45, 97-103, 470-519
Vulnerability Type: Unsafe third-party source and hook supply chain
Risk Level: Medium

Vulnerable Code Snippet

markdown
**Manual:**
```bash
git clone https://github.com/peterskoett/self-improving-agent.git ~/.openclaw/skills/self-improving-agent
text

```markdown
### Optional: Enable Hook

For automatic reminders at session start:

```bash
# Copy hook to OpenClaw hooks directory
cp -r hooks/openclaw ~/.openclaw/hooks/self-improvement

# Enable it
openclaw hooks enable self-improvement
text

```markdown
### Quick Setup (Claude Code / Codex)

Create `.claude/settings.json` in your project:

```json
{
  "hooks": {
    "UserPromptSubmit": [{
      "matcher": "",
      "hooks": [{
        "type": "command",
        "command": "./skills/self-improvement/scripts/activator.sh"
      }]
    }]
  }
}

This injects a learning evaluation reminder after each prompt (~50-100 tokens overhead).

text

```markdown
### Full Setup (With Error Detection)

```json
{
  "hooks": {
    "UserPromptSubmit": [{
      "matcher": "",
      "hooks": [{
        "type": "command",
        "command": "./skills/self-improvement/scripts/activator.sh"
      }]
    }],
    "PostToolUse": [{
      "matcher": "Bash",
      "hooks": [{
        "type": "command",
        "command": "./skills/self-improvement/scripts/error-detector.sh"
      }]
    }]
  }
}
text

### Technical Analysis

The manual installation procedure clones the current state of a mutable remote repository without pinning a reviewed commit, release tag, checksum, or cryptographic signature. The resulting files can then be installed as OpenClaw hooks or configured as command hooks that execute after user prompts and Bash tool activity.

This artifact contains only `SKILL.md`; the referenced `activator.sh`,
...[truncated 1748 chars]
Remediation
View remediation

Remediation Suggestions

  1. Pin installation to a specific reviewed commit hash or immutable, signed release.
  2. Publish expected SHA-256 checksums and verify downloaded content before installation.
  3. Require signed tags or release artifacts and document signature verification.
  4. Bundle all required hook scripts in the reviewed skill artifact so their exact contents can be audited.
  5. Display scripts to the user and require explicit approval before copying or enabling them.
  6. Keep automatic hooks disabled by default.
  7. Run hooks with minimal permissions, a restricted environment, and network access disabled unless explicitly required.
  8. Use absolute, verified script paths and ensure that repository contributors cannot silently replace the invoked files.
  9. Review and pin updates separately rather than automatically tracking the upstream default branch.
Vulnerability Patterns
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (3)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The description says to use the skill when a command fails, when a better approach is discovered, and to 'also review learnings before major tasks,' which are broad conditions that can overlap with normal day-to-day agent behavior. It does not provide clear boundaries or exclusion conditions for when the skill should not activate.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
88% confidence
Finding

The skill instructs the agent to create and populate persistent files under a user home workspace, and elsewhere encourages promotion of conversational learnings into durable memory files. In agents with automatic hooks or workspace injection, this can persist sensitive user content, errors, commands, or cross-session context beyond the original interaction, increasing the risk of privacy leakage and unintended retention.

Content

Scanner excerpt · SKILL.md (reported line 64)May include surrounding context.

└── FEATURE_REQUESTS.md

text

### Create Learning Files

```bash
mkdir -p ~/.openclaw/workspace/.learnings

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

Phrases like 'Can you also...' and 'Is there a way to...' are common everyday requests and are not specific to feature requests for this skill. Using them as activation signals risks unintended invocation during ordinary conversation.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.