T02 · Agent Memory Poisoning
- Location
SKILL.md:282- Finding
Untrusted Conversation Content Can Be Promoted into Persistent Agent Instructions
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is not deceptive or destructive, but it encourages broad persistent agent memory and optional hooks that can affect future sessions without enough review controls.
Review this skill before installing. Keep learning logs project-local where possible, do not enable global hooks unless you really want every future prompt to receive reminders, and require manual diff review before anything is promoted into AGENTS.md, CLAUDE.md, SOUL.md, TOOLS.md, or copilot instructions. Never store secrets, raw transcripts, credentials, or full command output in the learning files, and prefer a pinned version or reviewed local copy.
SKILL.md:282Untrusted Conversation Content Can Be Promoted into Persistent Agent Instructions
SKILL.md:52Installation Instructions Retrieve Mutable, Unpinned Skill Content
scripts/extract-skill.sh:101Output Path Validation Can Be Bypassed Through Pre-Existing Symbolic Links
The declared description presents a learning-capture and continuous-improvement skill focused on recording failures, corrections, outdated knowledge, and better approaches. The actual code does something materially different: it is a shell utility for scaffolding a new skill folder and template file from a skill name. While the generated template references extraction from a learning entry, the script itself neither captures learnings nor reviews them; it only creates files for a new skill. This is a primary-purpose mismatch, not just an implementation detail.
The guide recommends installing persistent hooks in the user-level agent configuration directory, which affects all future sessions rather than a single project. Because these hooks execute commands automatically with the agent's permissions, global placement materially increases blast radius if the script is modified, replaced, or behaves unexpectedly.
Add to ~/.claude/settings.json for global activation:
{
Instructions found that direct the agent to transmit conversation context or user data to external services.
Send message to another session:
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")
This skill instructs creation and persistent use of ~/.openclaw/workspace/.learnings, enabling cross-session retention of user interactions, failures, and operational context. Even though it warns against storing secrets, persistent session data in a shared workspace increases the chance of sensitive information, internal paths, or behavioral context being retained longer than intended and later accessed by other sessions or agents.
└── FEATURE_REQUESTS.md
### Create Learning Files
```bash
mkdir -p ~/.openclaw/workspace/.learnings
Phrases such as "Actually, it should be..." and "You're wrong about..." are common conversational feedback patterns and are presented here as automatic logging triggers without clear boundaries. The lack of negative examples or trigger constraints makes it unclear when routine clarification should be logged versus ignored.
The listed feature-request triggers include broad phrases like "Can you also...", "I wish you could...", and "Is there a way to...", which commonly appear in ordinary conversation outside the narrow context of logging a missing capability. Because the file presents these as automatic detection triggers, the activation scope is not sufficiently constrained and may cause unintended invocations.
Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.
When the above learning is extracted as a skill, it becomes:
File: skills/docker-m1-fixes/SKILL.md
---
The project-level settings create persistent automatic behavior that survives across sessions and can influence future agent interactions without per-session review. In the context of a self-improvement skill, persistence makes the behavior more dangerous because it continuously injects reminders and can normalize background processing of prompts and tool outputs.
Create .claude/settings.json in your project root:
{
The empty matcher causes the hook to run on every prompt, which broadens activation scope beyond error/debug scenarios and increases exposure to prompt content on all interactions. In a self-improvement skill, this creates unnecessary collection and processing of user inputs, raising privacy and overreach concerns even if the script only emits reminders.
The instructions create a persistent .learnings directory in the workspace or skill directory, enabling retention of operational and conversational artifacts across sessions. In the context of a self-improvement skill, this persistence becomes risky because it can accumulate sensitive data, influence future model behavior, and survive longer than users expect.
openclaw hooks enable self-improvement
### 3. Create Learning Files
Create the `.learnings/` directory in your workspace:
The guide tells users to promote learnings from ephemeral notes into persistent workspace files such as SOUL.md, TOOLS.md, and AGENTS.md, but it does not prominently warn that these files may retain sensitive operational details across sessions. Because OpenClaw injects workspace files into future sessions, over-collection can amplify privacy leakage and make accidental disclosure more likely.
The detection triggers are broad enough to fire on normal user corrections, tool errors, and vague 'knowledge gaps', which can cause excessive or unintended logging of conversational content into persistent memory. In a self-improvement skill, this increases the chance that sensitive prompts, mistakes, or context from unrelated tasks are captured without sufficient minimization or consent.
Conditions like "User provides information you didn't know" and "Unexpected output or behavior" depend on broad judgment calls and do not clearly define the threshold for invoking the skill. This ambiguity increases the chance of inconsistent or excessive activation.
No suspicious patterns detected.