Back to skill

Security audit

self-improving-agent

Security checks for vulnerabilities and agentic risk

Overview

The skill has a coherent learning-capture purpose, but it saves and promotes potentially sensitive conversation and error details into persistent agent files without enough redaction or approval safeguards.

Install only if you are comfortable with persistent learning files and optional hooks affecting future sessions. Keep it project-scoped where possible, do not enable global hooks casually, and redact tokens, credentials, private paths, customer data, request bodies, and full stack traces before writing learnings. Treat any promotion into AGENTS.md, SOUL.md, TOOLS.md, or MEMORY.md as a manual review step by a trusted operator, not an automatic result of repeated corrections.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T02 · Agent Memory Poisoning

Error
Location
hooks/openclaw/handler.js:11
Finding

Untrusted Corrections Can Be Promoted into Persistent Agent Control Files

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/error-detector.sh:9
Finding

Raw Tool Errors May Be Persisted Without Sensitive-Data Redaction

Content
View full analysis
A command error was detected. Consider logging this to .learnings/ERRORS.md if: - The error was unexpected or non-obvious - It required investigation to resolve - It might recur in similar contexts - The solution could benefit future sessions Use the self-improvement skill format: [ERR-YYYYMMDD-XXX] EOF fi ``` The associated error template requests complete error text and key environment information but does not require redaction of secrets or personal data. ### Technical Analysis The hook reads raw tool output from `CLAUDE_TOOL_OUTPUT` and encourages the agent to persist qualifying errors in `.learnings/ERRORS.md`. Command failures and stack traces commonly expose access tokens, authorization headers, connection strings, private filesystem paths, environment values, request payloads, and personal data. No redaction policy, secret-pattern filtering, output-size limit, retention policy, file-permission requirement, or source-control exclusion is implemented. Although the script does not automatically write the raw output, the combined hook and skill workflow explicitly encourages durable logging of the error information without a sanitization step. ### Attack Path 1. A command or integration fails and emits a credential, token, connection string, sensitive path, or personal data in its output. 2. `error-detector.sh` reads the raw output and detects one of its configured error patterns. 3. The hook prompts the agent to create an en ...[truncated 876 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (16)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的是一个“持续改进/经验记录”类能力:记录错误、纠正、能力缺口与最佳实践,并在任务前复盘历史经验。代码实际并未实现学习记录、错误归档、复盘检索或闭环分析等功能;它只是一个本地 shell 脚本,用来创建 skill 目录和 SKILL.md 模板文件。其主要能力是文件系统写入与模板生成,这与声明的核心用途明显不同,属于实质性目的不一致。

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The error template explicitly asks for full error text, parameters, and key environment information, which commonly contain API keys, internal URLs, usernames, file paths, stack traces, and customer data. Persisting these verbatim into markdown files materially increases the chance of secret leakage, lateral disclosure, and unintended retention of sensitive system context.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

The guide recommends adding hook configuration to ~/.claude/settings.json for global activation, which creates persistent execution of local shell scripts across sessions and repositories. This broadens the trust boundary from one project to the user’s entire agent environment, so any compromised or replaced referenced script could run automatically in many contexts.

Content

Scanner excerpt · references/hooks-setup.md (reported line 48)May include surrounding context.

Option 2: User-Level Configuration

Add to ~/.claude/settings.json for global activation:

json
{

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

Instructions found that direct the agent to transmit conversation context or user data to external services.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 177)May include surrounding context.

sessions_send

Send message to another session:

text
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill's natural-language instructions and examples are effectively constrained to Chinese throughout the file. Under the language/locale policy, forcing a specific language without user opt-in or documented regional justification is a policy concern.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill directs the agent to persist user corrections, capability requests, and long-term preferences/facts into durable memory files, creating a retention channel for conversation-derived data. Without explicit data classification and consent boundaries, this can accumulate sensitive user information and make later cross-task disclosure more likely.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill instructs persistent logging into workspace files but provides no guidance to avoid storing secrets, personal data, tokens, or sensitive operational context. Because these records are meant to be reused and reviewed later, accidental inclusion of sensitive information can create long-lived exposure in the repository or shared workspace.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

Detailed logging of raw errors, command context, parameters, and environment details creates a semantic exfiltration and retention risk even if no single field is labeled secret. These narratives often embed credentials, infrastructure topology, customer identifiers, and internal debugging data that can later be searched, shared, or committed.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

Recommending synchronization of learnings to MEMORY.md and sharing via sessions_send extends the exposure surface from local retention to cross-session transmission. If those learnings include conversation-derived or operationally sensitive details, this guidance can propagate data beyond its original context and increase the risk of inadvertent disclosure.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The template tells authors to 'Include trigger conditions' but does not require specific trigger phrases, scope boundaries, or negative examples. Because this is a reusable markdown template for manifests, it may propagate vague activation descriptions into downstream skills.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The minimal template uses the generic instruction 'What this skill does and when to use it' without requiring precise invocation constraints. This can lead authors to write broad descriptions that overlap with ordinary language and cause unintended activation.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · references/examples.md (reported line 301)May include surrounding context.

When the above learning is extracted as a skill, it becomes:

File: skills/docker-m1-fixes/SKILL.md

markdown
---

Session Persistence

Medium
Category
Rogue Agent
Confidence
82% confidence
Finding

Project-level hook configuration in .claude/settings.json establishes session-persistent behavior that automatically injects or executes logic during future agent interactions. While persistence is the intended feature, it is security-relevant because it survives beyond the immediate task and can continue affecting prompts and tool flows without repeated user review.

Content

Scanner excerpt · references/hooks-setup.md (reported line 15)May include surrounding context.

Option 1: Project-Level Configuration

Create .claude/settings.json in your project root:

json
{

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The document’s security section states that the scripts 'only output text' and 'don’t modify files or run commands', but the setup earlier explicitly configures them as command hooks that execute shell scripts. This mismatch is dangerous because it can cause users to grant trust and install auto-executed hooks under false assumptions, reducing scrutiny of code that runs on every prompt or tool event.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
87% confidence
Finding

The guide instructs users to create persistent learning storage in the workspace or skill directory, which increases the chance that sensitive operational details, user data, or credentials may be retained and later re-injected into future sessions. In this skill's context, persistent memory is a core mechanism, so omissions around data classification, retention, and cleanup make the risk more concrete.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 57)May include surrounding context.

openclaw hooks enable self-improvement

text

### 3. Create Learning Files

Create the `.learnings/` directory in your workspace:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The document encourages storing learnings in shared workspace files and using cross-session communication, but it does not warn against placing secrets, credentials, personal data, or sensitive internal details into those locations. In a system built around prompt injection from workspace files and transcript visibility across sessions, this can cause unintended propagation and retention of sensitive data beyond the original task context.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.