T02 · Agent Memory Poisoning
- Location
hooks/openclaw/handler.js:11- Finding
Persistent agent-context poisoning through untrusted learning promotion
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is mostly coherent, but it can persist broad agent-behavior changes from routine conversation into future sessions without strong review controls.
Review carefully before installing. Use project-local .learnings storage, keep hooks opt-in, avoid the command-output detector unless needed, and require explicit human approval before anything is promoted into CLAUDE.md, AGENTS.md, SOUL.md, TOOLS.md, or other automatically loaded agent context files.
hooks/openclaw/handler.js:11Persistent agent-context poisoning through untrusted learning promotion
hooks/openclaw/handler.js:28Shipped JavaScript hook injects instructions into sub-agent sessions despite source-level exclusion
scripts/extract-skill.sh:107Relative output containment can be bypassed through symbolic links
The declared description says this skill is for capturing learnings, errors, corrections, API/tool failures, and updated knowledge for continuous improvement. The supplied code does not capture, store, review, or manage learning entries. Instead, it is a helper script that creates a new skill scaffold under a skills directory and writes a templated SKILL.md file. While it references extracting a skill from a learning entry conceptually, its concrete behavior is skill generation, not learning capture or review. This is a materially different primary purpose, so the description does not accurately represent the code.
Skill reads from agent configuration directories (.claude/, .codex/, .gemini/). These directories may contain API keys, personal settings, and other credentials that the skill has no legitimate need to access.
Add to ~/.claude/settings.json for global activation:
{
Instructions found that direct the agent to transmit conversation context or user data to external services.
Send message to another session:
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")
The skill encourages persistent storage under a workspace home directory, which can retain error details, corrections, and feature requests across sessions. Persistence itself is not malicious, but it increases the exposure window for sensitive operational context, especially in shared machines, synced home directories, or repos where .learnings/ may later be committed or inspected by other tools.
└── FEATURE_REQUESTS.md
### Create Learning Files
```bash
mkdir -p ~/.openclaw/workspace/.learnings
The skill introduces cross-session transcript access and messaging capabilities, which can expose data from other sessions beyond the immediate task context. Even though the text advises using these only in trusted environments and with explicit user consent, normalizing access to session history increases the risk of unintended disclosure of sensitive prompts, outputs, or file paths.
Phrases such as "No, that's not right..." and "Actually, it should be..." are common in everyday back-and-forth and are listed as automatic logging triggers. The skill does not provide exclusion conditions or boundaries to distinguish casual clarification from situations that should invoke the self-improvement workflow.
The feature-request detection triggers include common conversational phrases like "Can you also...", "I wish you could...", and "Is there a way to..." that frequently occur in ordinary chats outside the narrow context of logging learnings. Because the document presents these as activation cues without negative examples or tighter constraints, the trigger scope is overly broad for a markdown skill description.
The skill recommends hook scripts that inspect every user prompt and potentially command output, creating a broad passive surveillance mechanism over sensitive developer inputs and tool results. This expands collection far beyond manual learning capture and can accidentally capture secrets, internal paths, proprietary code fragments, or authentication failures if the scripts are misconfigured or too permissive.
The stated purpose is to capture learnings, errors, and corrections for continuous improvement, but this section adds a separate capability to generate new skills via extract-skill.sh and create skills/<skill-name>/SKILL.md. Creating new reusable skills is a materially different capability from maintaining local learning logs and is not justified by the manifest's described scope.
This markdown template tells authors to 'Include trigger conditions' but does not require those triggers to be specific, bounded, or accompanied by exclusions. Because this file is itself a template for future skill manifests, it can propagate vague activation language into derived skills.
The minimal template asks for 'What this skill does and when to use it' but does not require explicit trigger scope, constraints, or exclusion conditions. For manifest-style descriptions, this can produce overly broad invocation criteria.
Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.
When the above learning is extracted as a skill, it becomes:
File: skills/docker-m1-fixes/SKILL.md
---
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
Create .claude/settings.json in your project root:
{
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
openclaw hooks enable self-improvement
### 3. Create Learning Files
Create the `.learnings/` directory in your workspace:
The 'Detection Triggers' section lists generic conditions such as 'Knowledge gaps', 'API errors', and user corrections without clearly defining activation boundaries or exclusions. These phrases overlap with common situations in many sessions, so the skill's invocation conditions are not specific enough to avoid accidental triggering.
The section says to apply the skill whenever the agent discovers something non-obvious, corrects itself, learns conventions, hits errors, or finds better approaches. These conditions are conceptually broad and subjective, making it unclear when the skill should activate versus when ordinary task execution should continue without logging.
The Quick Reference table uses generic placeholders '[Trigger 1]' and '[Trigger 2]' without clarifying that triggers must avoid broad everyday language. In a template, this omission can lead authors to supply ambiguous activation phrases that overlap with normal conversation.
No suspicious patterns detected.