Back to skill

Security audit

xiaowei-skill

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly coherent, but it can persist broad agent-behavior changes from routine conversation into future sessions without strong review controls.

Review carefully before installing. Use project-local .learnings storage, keep hooks opt-in, avoid the command-output detector unless needed, and require explicit human approval before anything is promoted into CLAUDE.md, AGENTS.md, SOUL.md, TOOLS.md, or other automatically loaded agent context files.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T02 · Agent Memory Poisoning

Error
Location
hooks/openclaw/handler.js:11
Finding

Persistent agent-context poisoning through untrusted learning promotion

Content
View full analysis
Remediation
View remediation

T01 · Skill Instruction Hijacking

Warning
Location
hooks/openclaw/handler.js:28
Finding

Shipped JavaScript hook injects instructions into sub-agent sessions despite source-level exclusion

Content
View full analysis
{ // Safety checks for event structure if (!event || typeof event !== 'object') { return; } // Only handle agent:bootstrap events if (event.type !== 'agent' || event.action !== 'bootstrap') { return; } // Safety check for context if (!event.context || typeof event.context !== 'object') { return; } // Inject the reminder as a virtual bootstrap file // Check that bootstrapFiles is an array before pushing if (Array.isArray(event.context.bootstrapFiles)) { event.context.bootstrapFiles.push({ path: 'SELF_IMPROVEMENT_REMINDER.md', content: REMINDER_CONTENT, virtual: true, }); } }; ``` The TypeScript source contains an exclusion that is absent from the shipped JavaScript: ```typescript // Skip sub-agent sessions to avoid bootstrap issues // Sub-agents have sessionKey patterns like "agent:main:subagent:..." const sessionKey = event.sessionKey || ''; if (sessionKey.includes(':subagent:')) { return; } ``` ### Technical Analysis OpenClaw commonly executes the JavaScript handler at runtime. Unlike `handler.ts`, `handler.js` does not inspect `event.sessionKey` and therefore injects `SELF_IMPROVEMENT_REMINDER.md` into sub-agent bootstrap contexts as well as the main agent. The discrepancy indicates that the intended scope restriction was added to the TypeScript source but was not propagated to the executable artifact. The injected content changes the sub-agent's active goals by directing it to capture and promote learnings, even when the delegated sub-agent task does not require this behavior. Because sub-agents may be used to process untrusted repositories, exte ...[truncated 1293 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/extract-skill.sh:107
Finding

Relative output containment can be bypassed through symbolic links

Content
View full analysis
"$SKILL_PATH/SKILL.md" << TEMPLATE ``` ### Technical Analysis Lexical validation does not guarantee that a relative path remains beneath the current workspace. Any existing component of `SKILLS_DIR` can be a symbolic link to a directory outside the workspace. For example, if `./skills` points to `/tmp/external`, the default output path `./skills/example/SKILL.md` resolves to `/tmp/external/example/SKILL.md`. The script nevertheless considers it a path under the current directory because it validates only the supplied string. There is also a time-of-check/time-of-use window between the existence check, directory creation, and file redirection. A local process capable of changing path components during execution may redirect the final write after validation. The script runs with the invoking agent's permissions, so the escaped write can reach any filesystem location writable by that account. ### Attack Path 1. An attacker gains the ability to prepare or modify files in a workspace used by the agent, such as through a malicious repository checkout. 2. The attacker creates `skills` or another accepted ou ...[truncated 1360 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (17)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description says this skill is for capturing learnings, errors, corrections, API/tool failures, and updated knowledge for continuous improvement. The supplied code does not capture, store, review, or manage learning entries. Instead, it is a helper script that creates a new skill scaffold under a skills directory and writes a templated SKILL.md file. While it references extracting a skill from a learning entry conceptually, its concrete behavior is skill generation, not learning capture or review. This is a materially different primary purpose, so the description does not accurately represent the code.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Skill reads from agent configuration directories (.claude/, .codex/, .gemini/). These directories may contain API keys, personal settings, and other credentials that the skill has no legitimate need to access.

Content

Scanner excerpt · references/hooks-setup.md (reported line 48)May include surrounding context.

Option 2: User-Level Configuration

Add to ~/.claude/settings.json for global activation:

json
{

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

Instructions found that direct the agent to transmit conversation context or user data to external services.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 181)May include surrounding context.

sessions_send

Send message to another session:

text
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")

Session Persistence

Medium
Category
Rogue Agent
Confidence
88% confidence
Finding

The skill encourages persistent storage under a workspace home directory, which can retain error details, corrections, and feature requests across sessions. Persistence itself is not malicious, but it increases the exposure window for sensitive operational context, especially in shared machines, synced home directories, or repos where .learnings/ may later be committed or inspected by other tools.

Content

Scanner excerpt · SKILL.md (reported line 81)May include surrounding context.

└── FEATURE_REQUESTS.md

text

### Create Learning Files

```bash
mkdir -p ~/.openclaw/workspace/.learnings

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The skill introduces cross-session transcript access and messaging capabilities, which can expose data from other sessions beyond the immediate task context. Even though the text advises using these only in trusted environments and with explicit user consent, normalizing access to session history increases the risk of unintended disclosure of sensitive prompts, outputs, or file paths.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

Phrases such as "No, that's not right..." and "Actually, it should be..." are common in everyday back-and-forth and are listed as automatic logging triggers. The skill does not provide exclusion conditions or boundaries to distinguish casual clarification from situations that should invoke the self-improvement workflow.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The feature-request detection triggers include common conversational phrases like "Can you also...", "I wish you could...", and "Is there a way to..." that frequently occur in ordinary chats outside the narrow context of logging learnings. Because the document presents these as activation cues without negative examples or tighter constraints, the trigger scope is overly broad for a markdown skill description.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill recommends hook scripts that inspect every user prompt and potentially command output, creating a broad passive surveillance mechanism over sensitive developer inputs and tool results. This expands collection far beyond manual learning capture and can accidentally capture secrets, internal paths, proprietary code fragments, or authentication failures if the scripts are misconfigured or too permissive.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The stated purpose is to capture learnings, errors, and corrections for continuous improvement, but this section adds a separate capability to generate new skills via extract-skill.sh and create skills/<skill-name>/SKILL.md. Creating new reusable skills is a materially different capability from maintaining local learning logs and is not justified by the manifest's described scope.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This markdown template tells authors to 'Include trigger conditions' but does not require those triggers to be specific, bounded, or accompanied by exclusions. Because this file is itself a template for future skill manifests, it can propagate vague activation language into derived skills.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The minimal template asks for 'What this skill does and when to use it' but does not require explicit trigger scope, constraints, or exclusion conditions. For manifest-style descriptions, this can produce overly broad invocation criteria.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · references/examples.md (reported line 301)May include surrounding context.

When the above learning is extracted as a skill, it becomes:

File: skills/docker-m1-fixes/SKILL.md

markdown
---

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · references/hooks-setup.md (reported line 15)May include surrounding context.

Option 1: Project-Level Configuration

Create .claude/settings.json in your project root:

json
{

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 57)May include surrounding context.

openclaw hooks enable self-improvement

text

### 3. Create Learning Files

Create the `.learnings/` directory in your workspace:

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The 'Detection Triggers' section lists generic conditions such as 'Knowledge gaps', 'API errors', and user corrections without clearly defining activation boundaries or exclusions. These phrases overlap with common situations in many sessions, so the skill's invocation conditions are not specific enough to avoid accidental triggering.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

The section says to apply the skill whenever the agent discovers something non-obvious, corrects itself, learns conventions, hits errors, or finds better approaches. These conditions are conceptually broad and subjective, making it unclear when the skill should activate versus when ordinary task execution should continue without logging.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The Quick Reference table uses generic placeholders '[Trigger 1]' and '[Trigger 2]' without clarifying that triggers must avoid broad everyday language. In a template, this omission can lead authors to supply ambiguous activation phrases that overlap with normal conversation.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.