Back to skill

Security audit

Self Improving Agent

Security checks for vulnerabilities and agentic risk

Overview

This skill’s memory features are mostly disclosed, but its generated-skill writer can persist unreviewed agent instructions and write outside the intended directory.

Install only if you want cross-session memory and generated skills. Avoid logging secrets, credentials, private code, sensitive prompts, or internal paths. Manually review generated SKILL.md files before allowing them to load or auto-trigger, and fix the skill-name path validation issue before using it in shared or untrusted workflows.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
generate_skill.py:49
Finding

Arbitrary File Placement and Overwrite Through Unsanitized Skill Name

Content
View full analysis
Remediation
View remediation

T02 · Agent Memory Poisoning

Warning
Location
generate_skill.py:52
Finding

Persistent Agent Instruction Injection Through Unescaped Generated Skill Content

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (18)

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The declared description presents a broad self-improving agent with multiple learning layers and behavioral memory. The supplied code only covers one narrow slice: proactive skill generation/storage. It writes generated skill files, maintains a registry, and supports listing/searching those records. There is no logic for observing mistakes, storing corrections, tracking successes beyond a default counter field, remembering preferences, or proactively checking future actions to prevent repeated errors. Because the description claims substantially broader autonomous learning behavior than the code actually provides, this is a material description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The declared description promises a multi-layer self-improving agent with learning from mistakes and successes, proactive checks, preference memory, and automatic skill generation. The supplied code only implements a narrow error-logging function: it parses CLI arguments and appends error records to a local JSONL file. While error capture is consistent with one small part of the description ('learn from mistakes'), the overall declared purpose materially overstates the behavior present in this code chunk. Therefore this chunk does not accurately represent the full declared functionality.

Content

No source excerpt is available for this finding.

YARA rule 'agent_skill_mcp_tool_poisoning_metadata': MCP/tool metadata poisoning indicators in tool schemas or skill manifests [agent_skills]

High
Category
YARA Match
Confidence
80% confidence
Finding

YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

Content

Scanner excerpt · SKILL.md (reported line 3)May include surrounding context.

md
---
name: self-improving-agent
description: "AI Self-Improving Agent v2 - Learn from mistakes, corrections, and successes. Three-layer system: passive capture + proactive check + proactive skill generation (inspired by Hermes Agent). Avoid repeating the same errors, remember user preferences, and auto-generate reusable Skills from successful task completions."
version: 2.0.0
trigger: "improve myself, learn from mistake, remember this, auto-skill, self-improving"
tags: ["memory", "learning", "skill-generation", "ai-agent", "productivity"]
license: MIT-0
author: 老二
platform: openclaw
---

# Self-Improving Agent v2

让AI�

YARA rule 'agent_skill_mcp_tool_poisoning_metadata': MCP/tool metadata poisoning indicators in tool schemas or skill manifests [agent_skills]

High
Category
YARA Match
Confidence
80% confidence
Finding

YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

Content

Scanner excerpt · skill.json (reported line 3)May include surrounding context.

json
{
  "name": "self-improving-agent",
  "description": "AI Self-Improving Agent v2 - Learn from mistakes, corrections, and successes. Three-layer system: passive capture + proactive check + proactive skill generation. Inspired by Hermes Agent's 'do-learn-fix' loop.",
  "version": "2.0.0",
  "trigger": [
    "improve myself",
    "learn from mistake", 
    "remember this",
    "auto-skill generation",
    "self-improving",
    "remember my preference",
    "don't repeat same error"
  ],
  "tags": [
    "memory",
    "learning",
    "skill-generation", 
    "ai-agent",
    "productivity",
    "automation"
  ],
  "license": "MIT-0",

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The README explicitly promotes persistent storage of errors, corrections, and best practices, but does not warn users that this memory may contain sensitive prompts, commands, file paths, proprietary code details, or personal preferences. In a self-improving agent, retained interaction history can accumulate sensitive data over time, increasing privacy and data-governance risk if stored insecurely or without informed consent.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
89% confidence
Finding

The skill is designed around session persistence by creating durable memory and generated-skill directories under the user's home directory. In this context, persistence is security-relevant because the agent stores behavioral history and can generate reusable skills, which may preserve sensitive data or unsafe patterns across sessions if not reviewed and constrained.

Content

Scanner excerpt · README.md (reported line 21)May include surrounding context.

或直接使用 skills-related 目录

创建必要目录

mkdir -p ~/.openclaw/memory/self-improving mkdir -p ~/.openclaw/skills-generated

text

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
96% confidence
Finding

The skill describes persistent file reads/writes and environment-path usage but does not declare any tool scope or allowed-tools boundary. That omission weakens reviewability and can let a host agent grant broader capabilities than users expect, especially for a skill that stores data and generates new artifacts.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill explicitly aims to remember mistakes, corrections, successes, and user preferences across sessions, creating a durable repository of potentially sensitive natural-language data. In context, this is especially risky because the stored content is unconstrained free text, which may include credentials, internal paths, personal preferences, or proprietary workflow details.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrases are broad enough to match ordinary conversation such as 'remember this' or 'learn from mistake,' which can activate the skill unintentionally. In this skill's context, accidental activation is more dangerous because activation can lead to persistent logging and skill-generation behavior without a deliberate, scoped request.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The main descriptive content switches into Chinese without indicating that language is optional or user-selectable. This can violate language/locale policy expectations because the skill appears to impose a specific language rather than offering localization or documenting a region-specific purpose.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The documentation advertises automatic capture of errors, corrections, and best practices without a clear warning that these interactions will be stored persistently. That creates consent and privacy risk because users may reveal sensitive information in corrections or task context that gets retained beyond the session.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The passive-capture layer directs the agent to log user corrections and other interaction data automatically into persistent files. Automatic capture increases the chance that sensitive context is stored without review, and this risk is amplified by the skill's broad triggers and self-improvement framing.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
91% confidence
Finding

The skill intentionally creates persistent directories under the user's home path for memory and generated skills, enabling cross-session state retention. Persistence is not inherently unsafe, but here it materially increases privacy, poisoning, and unauthorized reuse risks because the retained data influences future agent behavior.

Content

Scanner excerpt · SKILL.md (reported line 50)May include surrounding context.

bash
# Install
mkdir -p ~/.openclaw/memory/self-improving
mkdir -p ~/.openclaw/skills-generated

# Log an error

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The persistent files for corrections, best practices, and index data are defined structurally but without meaningful content restrictions, retention limits, or access controls. That invites overcollection and later leakage of operational or personal information through memory files and generated skills.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code creates directories and writes both a generated SKILL.md file and a registry JSON file, including replacing an existing registry entry when names match. Although it prints success messages afterward, there is no prior warning, confirmation, or inline disclosure that running the command will modify persistent files under the OpenClaw home directory.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The description advertises learning and skill generation but does not clearly warn users that their corrections, errors, preferences, or task outputs may be persistently captured and reused to create new skills. Missing disclosure is dangerous here because the agent's core function involves state retention and self-modification, which materially changes privacy, consent, and integrity risk.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The trigger phrases are generic conversational language such as 'remember this' and 'improve myself', which can cause the skill to activate during ordinary user interactions without clear consent. In a self-improving, memory-writing, and skill-generating agent, accidental activation increases the chance of unintended persistence, behavioral changes, or autonomous follow-on actions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The natural-language instructions and usage descriptions are presented in Chinese, which can impose a language constraint on users. There is no indication that the skill offers multilingual support, user opt-in, or that the locale limitation is required for a region-specific purpose.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.