Back to skill

Security audit

Self Improving Agent

Security checks for vulnerabilities and agentic risk

Overview

This skill is not malicious, but it asks the agent to persist user-derived lessons into future agent instructions with broad automatic triggers and too little review control.

Review this before installing if you use agents on sensitive work. Do not enable global hooks unless you want this behavior in every project, and require explicit review before anything from .learnings is written into AGENTS.md, SOUL.md, TOOLS.md, MEMORY.md, or a new skill.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:154
Finding

Unvalidated Promotion of User-Derived Learnings into Persistent Agent Instructions

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (24)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description is about maintaining a closed-loop learning process: capturing mistakes, corrections, capability gaps, and lessons learned for future review. The supplied code does not implement any learning log, classification, review, memory integration, or recurrence prevention workflow. Instead, it is a helper script for creating a new skill scaffold on disk from a skill name. This is materially closer to a skill-creation utility, which the declared description explicitly says it is NOT for. The code’s primary purpose, triggers, and resource access (writing directories/files under ./skills) differ substantially from the declared purpose.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Instructing users to modify ~/.claude/settings.json places persistent executable hook configuration in the agent’s user-level config directory. That is dangerous because it establishes long-lived behavior across repositories and trust zones, and if abused or misconfigured can silently affect future sessions.

Content

Scanner excerpt · references/hooks-setup.md (reported line 48)May include surrounding context.

Option 2: User-Level Configuration

Add to ~/.claude/settings.json for global activation:

json
{

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

Instructions found that direct the agent to transmit conversation context or user data to external services.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 177)May include surrounding context.

sessions_send

Send message to another session:

text
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The script explicitly implements skill creation and scaffolding, which exceeds the stated scope of a self-improvement skill focused on recording, reviewing, and integrating learnings. This scope expansion is dangerous because it grants file-creation capabilities that can be repurposed to introduce new executable artifacts and persistence mechanisms under the guise of a benign learning workflow.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

This code creates directories and writes a new SKILL.md file, moving beyond passive learning capture into active repository modification. In the context of an agent skill, unauthorized write/scaffolding capability increases the risk of privilege creep, hidden persistence, and creation of follow-on artifacts that may later be executed or trusted by operators.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The README promotes automatic error logging, learning capture, memory integration, and rule formation, but does not warn users that these actions may persist prompts, corrections, failures, or other potentially sensitive data. This creates a privacy and security risk because users may unknowingly allow durable storage of confidential context that later influences agent behavior.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
90% confidence
Finding

The design explicitly calls for promoting patterns to long-term memory and forming hardened rules to prevent recurrence, which is a form of session persistence that can carry information and behavior modifications across tasks. Without strict controls, this can preserve sensitive data, encode mistaken assumptions permanently, or let adversarial/user-supplied corrections shape future behavior beyond the original session.

Content

Scanner excerpt · README.md (reported line 11)May include surrounding context.

md
- 🎓 **Learning Capture**: Record corrections and discoveries
- 🔄 **Periodic Review**: Regular retrospectives to consolidate learnings
- 🧠 **Memory Integration**: Promote important patterns to long-term memory
- 🛡️ **Rule Formation**: Create hardened rules to prevent recurring mistakes

## How It Works

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The activation conditions are broad enough that the skill may trigger on many ordinary failures, corrections, or capability gaps without a clear user opt-in boundary. In a skill that performs persistent learning and memory updates, ambiguous triggering increases the chance of over-collection, unintended retention, and rule changes based on low-quality or sensitive interactions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest description defines when to use the skill in a fixed mix of Chinese and English, which can effectively impose a language/locale expectation on users or downstream agents. The policy allows locale constraints only when they are justified or when the user is offered a language choice, neither of which is stated here.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

These instructions explicitly tell the agent to consolidate learnings into persistent files such as MEMORY.md and other long-term rule files. That creates a real data-retention risk because user-provided content, mistakes, and context may be stored in plain language without minimization, consent, retention limits, or sensitivity filtering.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The suggested MEMORY.md structure explicitly includes sections for 'user preferences, habits, important information,' encouraging persistent storage of personal and potentially sensitive data. In this context the danger is elevated because the design normalizes long-term accumulation of user-specific data in plaintext workspace files, increasing privacy, insider-access, and accidental disclosure risk.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The trigger list includes broad everyday phrases like correction requests and generic feature-request language, which can cause the skill to activate in situations where the user did not intend durable logging or memory updates. In this skill's context, unintended invocation is more dangerous because activation can lead to persistence of conversational content into local memory files, creating privacy and over-collection risks.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · references/examples.md (reported line 301)May include surrounding context.

When the above learning is extracted as a skill, it becomes:

File: skills/docker-m1-fixes/SKILL.md

markdown
---

Session Persistence

Medium
Category
Rogue Agent
Confidence
80% confidence
Finding

Creating .claude/settings.json in the project root introduces persistent session behavior that can survive across runs and collaborators if committed or reused. In this context, persistence makes broad automatic hook execution more dangerous because it may continue operating after the original need or review context is gone.

Content

Scanner excerpt · references/hooks-setup.md (reported line 15)May include surrounding context.

Option 1: Project-Level Configuration

Create .claude/settings.json in your project root:

json
{

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

An empty matcher makes the hook fire on every prompt, creating broad automatic execution during normal agent use. In this skill context, that increases exposure to prompt-triggered behavior, unnecessary data processing, and accidental capture of sensitive or irrelevant session content.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

User-level configuration in ~/.claude/settings.json enables the hook across all projects and sessions, greatly expanding scope and persistence. In a self-improvement skill, that means lessons, prompts, and workflow metadata may be processed globally, increasing the blast radius if the scripts are flawed or later modified.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The minimal setup still uses an empty matcher, so it remains effectively always-on despite being presented as lower overhead. That framing may cause users to underestimate the security and privacy implications of broad automatic triggering.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The Codex example also uses an empty matcher, causing the hook to overlap with routine prompt traffic rather than specific learning events. This broad scope is risky because it normalizes automatic execution and potential collection in contexts where the skill is not needed.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The document’s security section claims the scripts only output text and do not run commands, but the hook configuration explicitly uses command hooks and also references an extraction script that creates scaffolding. This mismatch can mislead users about the trust boundary and execution model, causing them to enable hooks with more privilege and side effects than they understand.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 57)May include surrounding context.

openclaw hooks enable self-improvement

text

### 3. Create Learning Files

Create the `.learnings/` directory in your workspace:

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The "Detection Triggers" section uses broad phrases such as "Knowledge gaps" and generic events like "API errors" and "Command failures" without defining scope, exclusions, or thresholds. In a markdown skill/integration document, this can cause the skill or logging workflow to activate in many ordinary situations rather than only in clearly intended cases.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The trigger example "User corrections ("No, that's wrong...")" is a common conversational phrase that could occur in many contexts, but the document does not constrain when it should invoke the learning behavior. Without boundaries, this phrasing is too broad for reliable activation.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The ability to scaffold new skills is not justified by the declared purpose of a learning-loop skill and broadens the agent's effective authority. Even though the script validates relative paths and blocks '..', it still enables creation of new agent-facing content and guidance, which can be abused to plant misleading or executable components in the workspace.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The manual trigger examples include a Chinese-language phrase alongside English, but the README does not explain language support, whether multilingual triggers are optional, or how locale selection works. This can create an undocumented language/locale behavior rather than an explicit user opt-in.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.