T02 · Agent Memory Poisoning
- Location
SKILL.md:154- Finding
Unvalidated Promotion of User-Derived Learnings into Persistent Agent Instructions
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill is not malicious, but it asks the agent to persist user-derived lessons into future agent instructions with broad automatic triggers and too little review control.
Review this before installing if you use agents on sensitive work. Do not enable global hooks unless you want this behavior in every project, and require explicit review before anything from .learnings is written into AGENTS.md, SOUL.md, TOOLS.md, MEMORY.md, or a new skill.
SKILL.md:154Unvalidated Promotion of User-Derived Learnings into Persistent Agent Instructions
The declared description is about maintaining a closed-loop learning process: capturing mistakes, corrections, capability gaps, and lessons learned for future review. The supplied code does not implement any learning log, classification, review, memory integration, or recurrence prevention workflow. Instead, it is a helper script for creating a new skill scaffold on disk from a skill name. This is materially closer to a skill-creation utility, which the declared description explicitly says it is NOT for. The code’s primary purpose, triggers, and resource access (writing directories/files under ./skills) differ substantially from the declared purpose.
Instructing users to modify ~/.claude/settings.json places persistent executable hook configuration in the agent’s user-level config directory. That is dangerous because it establishes long-lived behavior across repositories and trust zones, and if abused or misconfigured can silently affect future sessions.
Add to ~/.claude/settings.json for global activation:
{
Instructions found that direct the agent to transmit conversation context or user data to external services.
Send message to another session:
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")
The script explicitly implements skill creation and scaffolding, which exceeds the stated scope of a self-improvement skill focused on recording, reviewing, and integrating learnings. This scope expansion is dangerous because it grants file-creation capabilities that can be repurposed to introduce new executable artifacts and persistence mechanisms under the guise of a benign learning workflow.
This code creates directories and writes a new SKILL.md file, moving beyond passive learning capture into active repository modification. In the context of an agent skill, unauthorized write/scaffolding capability increases the risk of privilege creep, hidden persistence, and creation of follow-on artifacts that may later be executed or trusted by operators.
The README promotes automatic error logging, learning capture, memory integration, and rule formation, but does not warn users that these actions may persist prompts, corrections, failures, or other potentially sensitive data. This creates a privacy and security risk because users may unknowingly allow durable storage of confidential context that later influences agent behavior.
The design explicitly calls for promoting patterns to long-term memory and forming hardened rules to prevent recurrence, which is a form of session persistence that can carry information and behavior modifications across tasks. Without strict controls, this can preserve sensitive data, encode mistaken assumptions permanently, or let adversarial/user-supplied corrections shape future behavior beyond the original session.
- 🎓 **Learning Capture**: Record corrections and discoveries
- 🔄 **Periodic Review**: Regular retrospectives to consolidate learnings
- 🧠 **Memory Integration**: Promote important patterns to long-term memory
- 🛡️ **Rule Formation**: Create hardened rules to prevent recurring mistakes
## How It Works
The activation conditions are broad enough that the skill may trigger on many ordinary failures, corrections, or capability gaps without a clear user opt-in boundary. In a skill that performs persistent learning and memory updates, ambiguous triggering increases the chance of over-collection, unintended retention, and rule changes based on low-quality or sensitive interactions.
The manifest description defines when to use the skill in a fixed mix of Chinese and English, which can effectively impose a language/locale expectation on users or downstream agents. The policy allows locale constraints only when they are justified or when the user is offered a language choice, neither of which is stated here.
These instructions explicitly tell the agent to consolidate learnings into persistent files such as MEMORY.md and other long-term rule files. That creates a real data-retention risk because user-provided content, mistakes, and context may be stored in plain language without minimization, consent, retention limits, or sensitivity filtering.
The suggested MEMORY.md structure explicitly includes sections for 'user preferences, habits, important information,' encouraging persistent storage of personal and potentially sensitive data. In this context the danger is elevated because the design normalizes long-term accumulation of user-specific data in plaintext workspace files, increasing privacy, insider-access, and accidental disclosure risk.
The trigger list includes broad everyday phrases like correction requests and generic feature-request language, which can cause the skill to activate in situations where the user did not intend durable logging or memory updates. In this skill's context, unintended invocation is more dangerous because activation can lead to persistence of conversational content into local memory files, creating privacy and over-collection risks.
Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.
When the above learning is extracted as a skill, it becomes:
File: skills/docker-m1-fixes/SKILL.md
---
Creating .claude/settings.json in the project root introduces persistent session behavior that can survive across runs and collaborators if committed or reused. In this context, persistence makes broad automatic hook execution more dangerous because it may continue operating after the original need or review context is gone.
Create .claude/settings.json in your project root:
{
An empty matcher makes the hook fire on every prompt, creating broad automatic execution during normal agent use. In this skill context, that increases exposure to prompt-triggered behavior, unnecessary data processing, and accidental capture of sensitive or irrelevant session content.
User-level configuration in ~/.claude/settings.json enables the hook across all projects and sessions, greatly expanding scope and persistence. In a self-improvement skill, that means lessons, prompts, and workflow metadata may be processed globally, increasing the blast radius if the scripts are flawed or later modified.
The minimal setup still uses an empty matcher, so it remains effectively always-on despite being presented as lower overhead. That framing may cause users to underestimate the security and privacy implications of broad automatic triggering.
The Codex example also uses an empty matcher, causing the hook to overlap with routine prompt traffic rather than specific learning events. This broad scope is risky because it normalizes automatic execution and potential collection in contexts where the skill is not needed.
The document’s security section claims the scripts only output text and do not run commands, but the hook configuration explicitly uses command hooks and also references an extraction script that creates scaffolding. This mismatch can mislead users about the trust boundary and execution model, causing them to enable hooks with more privilege and side effects than they understand.
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
openclaw hooks enable self-improvement
### 3. Create Learning Files
Create the `.learnings/` directory in your workspace:
The "Detection Triggers" section uses broad phrases such as "Knowledge gaps" and generic events like "API errors" and "Command failures" without defining scope, exclusions, or thresholds. In a markdown skill/integration document, this can cause the skill or logging workflow to activate in many ordinary situations rather than only in clearly intended cases.
The trigger example "User corrections ("No, that's wrong...")" is a common conversational phrase that could occur in many contexts, but the document does not constrain when it should invoke the learning behavior. Without boundaries, this phrasing is too broad for reliable activation.
The ability to scaffold new skills is not justified by the declared purpose of a learning-loop skill and broadens the agent's effective authority. Even though the script validates relative paths and blocks '..', it still enables creation of new agent-facing content and guidance, which can be abused to plant misleading or executable components in the workspace.
The manual trigger examples include a Chinese-language phrase alongside English, but the README does not explain language support, whether multilingual triggers are optional, or how locale selection works. This can create an undocumented language/locale behavior rather than an explicit user opt-in.
No suspicious patterns detected.