Back to skill

Security audit

Adityasagar

Security checks for vulnerabilities and agentic risk

Overview

The skill is a disclosed self-improvement logger, but it encourages saving conversation details and promoting them into persistent agent instructions without enough approval or redaction safeguards.

Install only if you are comfortable with a skill that creates long-lived learning notes and may encourage agents to update future instruction files. Keep logs local when possible, redact secrets and private context before saving, require explicit human review before anything is promoted into agent memory or instruction files, and avoid enabling cross-session sharing or hooks globally unless you need them.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:23
Finding

Conversation-Derived Content Can Poison Persistent Agent Instructions

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:23-26, SKILL.md:346-362, SKILL.md:448; hooks/openclaw/handler.js:8-24, 46-52; hooks/openclaw/handler.ts:10-25, 52-58
Vulnerability Type: Persistent agent memory poisoning
Risk Level: High

Vulnerable Code and Instructions

SKILL.md:23-26:

markdown
| Broadly applicable learning | Promote to `CLAUDE.md`, `AGENTS.md`, and/or `.github/copilot-instructions.md` |
| Workflow improvements | Promote to `AGENTS.md` (OpenClaw workspace) |
| Tool gotchas | Promote to `TOOLS.md` (OpenClaw workspace) |
| Behavioral patterns | Promote to `SOUL.md` (OpenClaw workspace) |

SKILL.md:346-362:

markdown
Promotion targets:
- `CLAUDE.md`
- `AGENTS.md`
- `.github/copilot-instructions.md`
- `SOUL.md` / `TOOLS.md` for OpenClaw workspace-level guidance when applicable

Write promoted rules as short prevention rules (what to do before/while coding),
not long incident write-ups.

SKILL.md:448:

markdown
7. **Promote aggressively** - if in doubt, add to CLAUDE.md or .github/copilot-instructions.md

hooks/openclaw/handler.js:8-24, 46-52:

javascript
const REMINDER_CONTENT = `
## Self-Improvement Reminder

After completing tasks, evaluate if any learnings should be captured:

**Log when:**
- User corrects you → \`.learnings/LEARNINGS.md\`
- Command/operation fails → \`.learnings/ERRORS.md\`
- User wants missing capability → \`.learnings/FEATURE_REQUESTS.md\`
- You discover your knowledge was wrong → \`.learnings/LEARNINGS.md\`
- You find a better approach → \`.learnings/LEARNINGS.md\`

**Promote when pattern is proven:**
- Behavioral patterns → \`SOUL.md\`
- Workflow improvements → \`AGENTS.md\`
- Tool gotchas → \`TOOLS.md\`

Keep entries simple: date, title, what happened, what to do differently.
`.trim();

if (Array.isArray(event.context.bootstrapFiles)) {
  event.context.bootstrapFiles.push({
  
...[truncated 2644 chars]
Remediation
View remediation

Remediation Suggestions

  1. Require explicit human approval before promoting any learning into an agent instruction file.
  2. Display the exact destination path and a complete diff before writing, and require a separate confirmation for the write.
  3. Store untrusted conversation-derived observations in a dedicated data file that is not loaded as authoritative instructions.
  4. Prohibit automatic promotion of content involving credentials, authentication, network access, command execution, security controls, instruction precedence, or requests to ignore existing policy.
  5. Add provenance metadata recording the source session, author, timestamp, evidence, reviewer, and approval decision.
  6. Require multiple independently verified observations before considering a learning for promotion; recurrence alone must not establish trust.
  7. Replace “promote aggressively” with a conservative rule that defaults to no promotion.
  8. Maintain version history and provide a straightforward rollback mechanism for all persistent instruction changes.
  9. Clearly separate factual project knowledge from behavioral or executable agent instructions.
  10. For the bootstrap hook, remind the agent that conversation content is untrusted and that promotion requires human review rather than merely encouraging promotion.

T08 · Insecure Dependencies

Warning
Location
SKILL.md:32
Finding

Installation Instructions Use Mutable, Unverified External Sources

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:32-44; references/openclaw-integration.md:31-40
Vulnerability Type: Unpinned external skill dependency
Risk Level: Medium

Vulnerable Code and Instructions

SKILL.md:32-44:

markdown
### Installation

**Via ClawdHub (recommended):**
```bash
clawdhub install self-improving-agent

Manual:

bash
git clone https://github.com/peterskoett/self-improving-agent.git ~/.openclaw/skills/self-improving-agent

Remade for openclaw from original repo : https://github.com/pskoett/pskoett-ai-skills - https://github.com/pskoett/pskoett-ai-skills/tree/main/skills/self-improvement

text

`references/openclaw-integration.md:31-40`:

```markdown
### 1. Install the Skill

```bash
clawdhub install self-improving-agent

Or copy manually:

bash
cp -r self-improving-agent ~/.openclaw/skills/
text

### Technical Analysis

The recommended registry installation does not specify an audited version, while the Git installation clones the repository’s mutable default branch without pinning a commit or tag. Neither procedure specifies a cryptographic checksum, signature, or expected package digest.

Consequently, the installed content at a later date may differ from the audited project snapshot. This is particularly sensitive because the package contains agent instructions, executable shell scripts, and an optional OpenClaw hook. A compromised publisher account, registry entry, repository, or upstream release process could therefore substitute altered content.

The reviewed snapshot does not contain remote payload retrieval followed by execution, and no evidence shows that the referenced upstream sources are currently malicious. The risk arises from the mutable and unverified installation process.

### Attack Path

1. An attacker compromises the registry package, publisher credentials, source repository, or release process.
2. The at
...[truncated 1062 chars]
Remediation
View remediation

Remediation Suggestions

  1. Pin registry installations to a specific audited version rather than implicitly installing the latest release.
  2. For Git installations, check out a documented immutable commit hash:
    bash
    git clone https://github.com/peterskoett/self-improving-agent.git
    cd self-improving-agent
    git checkout --detach <audited-commit-sha>
    
  3. Publish and verify a SHA-256 digest or signed release manifest for every distributed package.
  4. Use signed tags or releases and document signature-verification steps.
  5. Review the exact installed SKILL.md, hook handlers, and shell scripts before enabling hooks or invoking scripts.
  6. Treat updates as new security-sensitive installations: compare the update against the previously audited version and require approval.
  7. Configure the registry or installer to fail closed when integrity or signature verification cannot be completed.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (17)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description is about maintaining a repository of learnings, errors, and corrections to support continuous improvement and reviewing those learnings before major tasks. The supplied code does not capture, store, review, or analyze learnings. Instead, it is a file-generation helper that scaffolds a new skill directory and template markdown file based on a skill name. While the comments mention extraction from a learning entry, the implementation only creates a boilerplate skill structure and suggests manually updating the original learning entry afterward. This is a materially different primary purpose, so the description does not accurately represent the code chunk.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The learning-entry format asks for full context and related details, which strongly encourages copying conversation content, incident specifics, and file references into persistent logs. In a security context, this is dangerous because those notes can accumulate secrets, proprietary architecture details, or sensitive business logic in an easily searchable repository.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Skill reads from agent configuration directories (.claude/, .codex/, .gemini/). These directories may contain API keys, personal settings, and other credentials that the skill has no legitimate need to access.

Content

Scanner excerpt · references/hooks-setup.md (reported line 48)May include surrounding context.

Option 2: User-Level Configuration

Add to ~/.claude/settings.json for global activation:

json
{

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

Instructions found that direct the agent to transmit conversation context or user data to external services.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 177)May include surrounding context.

sessions_send

Send message to another session:

text
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")

Ssd 3

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The core workflow instructs the agent to persist conversation-derived corrections, feature requests, and contextual details into reusable files and memory stores. That creates a durable natural-language retention channel that can capture sensitive user information and make it available in later sessions beyond the original need.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
74% confidence
Finding

The skill explicitly establishes persistent storage under a workspace path, enabling retained memory across sessions. Persistence is not inherently malicious, but in this skill's context it materially increases the impact of the logging/privacy issues because retained data can be reloaded, searched, or shared later.

Content

Scanner excerpt · SKILL.md (reported line 64)May include surrounding context.

└── FEATURE_REQUESTS.md

text

### Create Learning Files

```bash
mkdir -p ~/.openclaw/workspace/.learnings

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The inter-session communication guidance introduces data-sharing capabilities unrelated to simple local learning capture, including transcript reading and sending learnings across sessions. In context, that broadens the trust boundary and can expose sensitive prompts, outputs, or operational context across agents without any privacy guardrails.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

Reading session history and sending learnings across sessions can disclose sensitive conversation contents, hidden instructions, credentials, or internal context to other agents or workstreams. The absence of warnings, redaction guidance, or approval gates makes this especially risky in a toolchain that persists or forwards transcripts.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The error logging template explicitly asks for actual error output, command context, inputs, and environment details, all of which commonly contain secrets, tokens, paths, personal data, or proprietary system information. Because the skill encourages persistent storage of this raw context, it creates a clear data-retention and inadvertent secret-disclosure risk.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The feature request template explicitly stores what the user wanted and why, which can encode confidential roadmap intent, internal business priorities, or personal context. Persisting that information in shared or long-lived files expands exposure well beyond the original conversation.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The promotion guidance encourages copying learned rules into persistent agent context files used in future sessions, increasing both longevity and surface area of any sensitive content captured from prior interactions. This makes accidental disclosure more likely because the information is elevated into always-loaded or broadly shared context.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The listed activation examples include very broad phrases such as "Can you also...", "Is there a way to...", and "Why can't you...", which commonly appear in ordinary conversation outside the skill's intended self-improvement context. The section does not provide exclusion conditions or constraints to distinguish routine requests from cases that should actually trigger logging.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest describes this skill as a mechanism for recording learnings and reviewing them for continuous improvement. The 'Automatic Skill Extraction' section adds a distinct capability: turning learnings into new skills via helper scripts and manual skill creation workflows, which goes beyond simple capture/review of learnings.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

Suggested prompts like "Log this to learnings" and "Create a skill from this solution" are short, generic commands that could collide with normal conversation and do not define when the skill should or should not activate. There are no guardrails or contextual qualifiers limiting these prompts to specific environments or workflows.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · references/examples.md (reported line 301)May include surrounding context.

When the above learning is extracted as a skill, it becomes:

File: skills/docker-m1-fixes/SKILL.md

markdown
---

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · references/hooks-setup.md (reported line 15)May include surrounding context.

Option 1: Project-Level Configuration

Create .claude/settings.json in your project root:

json
{

Session Persistence

Medium
Category
Rogue Agent
Confidence
81% confidence
Finding

The skill explicitly instructs users to create persistent learning storage under the workspace or skill directory, enabling information from prior sessions to be retained and later reinjected into agent context. In a prompt-injection-based system, persistent memory can accumulate sensitive data, user corrections, or adversarial instructions that survive resets and influence future sessions.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 57)May include surrounding context.

openclaw hooks enable self-improvement

text

### 3. Create Learning Files

Create the `.learnings/` directory in your workspace:

Static analysis

No suspicious patterns detected.