T02 · Agent Memory Poisoning
- Location
SKILL.md:24- Finding
Untrusted Learning Promotion Can Poison Persistent Agent Instructions
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is not clearly malicious, but it can persistently change agent memory and hooks while capturing session context too broadly.
Install only if you want a persistent agent-memory workflow. Keep it project-scoped, avoid global hooks, review every hook script and settings change, never store secrets or raw transcripts in .learnings, and require a visible diff plus explicit approval before promoting anything into CLAUDE.md, AGENTS.md, SOUL.md, TOOLS.md, or Copilot instructions.
SKILL.md:24Untrusted Learning Promotion Can Poison Persistent Agent Instructions
hooks/openclaw/handler.js:27Opt-In Hooks Persistently Inject Skill Instructions into Agent Context
SKILL.md:33Installation Instructions Use Unpinned Mutable Upstream Sources
The declared description is about capturing and using learnings for continuous improvement when errors, corrections, outdated knowledge, or tool failures occur. The supplied code does not implement learning capture, error/correction tracking, review of learnings, or any trigger-based continuous-improvement behavior. Instead, it is a helper utility for extracting/promoting a learning into a new skill scaffold by creating directories and a templated markdown file. While this may be adjacent to a broader learnings workflow, its primary purpose is materially different and includes undeclared filesystem-writing/scaffolding capabilities.
Cross-session transcript reading and message passing create a direct pathway for prior session data to be accessed or relayed into unrelated contexts. Without strict authorization and minimization rules, this can expose sensitive content from one user task to another session or agent that did not originally receive it.
The templates explicitly ask for 'full context,' inputs, parameters, environment details, and user context, which strongly encourages plaintext storage of secrets, tokens, internal paths, proprietary prompts, or personal data. This is dangerous because the notes become a secondary datastore that may bypass normal application privacy and secret-handling controls.
Directing users to modify ~/.claude/settings.json affects the agent's persistent global configuration, which is a sensitive control point. Changes there can silently influence future sessions and projects, so coupling it with automatic command hooks creates a durable execution mechanism that could be abused or accidentally left enabled far beyond the intended scope.
Add to ~/.claude/settings.json for global activation:
{
Instructions found that direct the agent to transmit conversation context or user data to external services.
Send message to another session:
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")
The manifest's broad activation conditions can cause the skill to trigger in many normal interactions, increasing the chance that conversation content is persistently logged or promoted when not necessary. In a system with automatic skill loading or hooks, overbroad triggering raises the attack surface for prompt-driven data retention and context pollution.
The skill encourages persistent logging and cross-session sharing of learnings, errors, and corrections, but provides no minimization, redaction, retention, or access-control guidance. As a result, user-provided details and task context may be stored and redistributed in plaintext beyond the original session, creating privacy and data leakage risks.
The instructions establish persistent local storage under a workspace directory, which creates durable retention of operational and conversational data. Persistence itself is not inherently unsafe, but in this skill it amplifies the privacy risks because the stored content may include detailed errors, user context, and learnings without retention or protection rules.
└── FEATURE_REQUESTS.md
### Create Learning Files
```bash
mkdir -p ~/.openclaw/workspace/.learnings
The trigger phrases are common conversational language such as corrections, wishes, and questions, so they can fire on ordinary chat and capture data without meaningful intent from the user. That makes accidental persistence and propagation of sensitive or irrelevant content more likely, especially when paired with hooks or automated review behavior.
The manifest describes a self-improvement skill focused on capturing learnings, errors, corrections, and reviewing/promoting them for continuous improvement. This section adds a distinct capability: extracting and scaffolding entirely new skills via helper scripts and manual creation workflows, which is not a direct or necessary part of logging learnings.
The template asks authors to include trigger conditions, but it does not require those conditions to be concrete, bounded, or testable. In an agent skill system, vague activation criteria can cause over-broad invocation, leading the agent to apply the skill in unintended contexts and potentially run inappropriate guidance or helper commands.
The minimal template allows a description of 'what this skill does and when to use it' without requiring precise constraints, exclusions, or verification conditions. Minimal templates are likely to be copied as-is, so this omission can propagate ambiguous activation logic across many skills.
The script-oriented template includes executable helpers but does not emphasize strict activation boundaries before those helpers are suggested or run. In practice, vague invocation criteria combined with scripts increases the risk that automation is applied in the wrong environment, against the wrong target, or without necessary preconditions.
Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.
When the above learning is extracted as a skill, it becomes:
File: skills/docker-m1-fixes/SKILL.md
---
Placing hook configuration in a project-local .claude/settings.json establishes persistent behavior for future sessions in that repository. While this is a legitimate feature, in this context it creates lasting automatic execution of local scripts, which can be risky if the repository is shared, the scripts change over time, or users do not realize the behavior persists.
Create .claude/settings.json in your project root:
{
An empty matcher on a UserPromptSubmit hook causes the script to run on every prompt, creating a very broad automatic execution surface. In this skill context, that means any session interaction can trigger local scripts, amplifying the impact of mistakes, malicious modifications to the script, or unintended data exposure through hook processing.
The user-level configuration installs the hook in ~/.claude/settings.json for global activation, so the script will run across projects and contexts without meaningful constraints. That broad persistence increases the blast radius of any compromised or buggy script and can affect unrelated repositories, sessions, or sensitive workflows.
The Codex example also uses an empty matcher, so the hook's invocation conditions are effectively unconstrained. In a tool that can access project context and run with user permissions, unconditional hook execution unnecessarily expands exposure and makes accidental or malicious behavior more likely to trigger.
The document's security section asserts that the hook scripts only output text and do not run commands, but the same file configures them as command hooks and instructs users to execute an extraction script directly. That inconsistency can mislead users into underestimating the trust and execution risk of these scripts, increasing the chance they enable arbitrary local code execution with their agent's permissions.
The instructions explicitly create a persistent .learnings/ store in the workspace or skill directory, enabling information from one interaction to survive into later sessions. In the context of a self-improvement skill, this persistence increases the blast radius of any accidental capture of sensitive data or any prompt-injected content that gets written as a 'learning.'
openclaw hooks enable self-improvement
### 3. Create Learning Files
Create the `.learnings/` directory in your workspace:
The guide encourages storing learnings in persistent files and promoting them into workspace-wide prompt files, then sharing them across sessions, but it provides no warnings or controls for sensitive data. This creates a realistic risk that secrets, personal data, internal prompts, or incident details entered during failures and corrections will be retained and propagated beyond the original context.
The trigger definitions are broad enough that routine failures, vague user corrections, or normal model uncertainty could cause the skill to activate and write persistent records without clear user intent. In a self-improvement skill, this can lead to excessive or inaccurate logging, and can amplify prompt-injected or attacker-induced events into durable memory.
Comments and usage indicate this tool 'creates a new skill from a learning entry,' giving the skill the capability to generate new project artifacts. For a skill whose stated purpose is capturing learnings, errors, and corrections, automatic creation of new skill directories and markdown manifests is a separate lifecycle-management capability not obviously required by that purpose.
The manifest describes a skill focused on recording and reviewing learnings, corrections, and failures to support self-improvement. This script instead scaffolds entirely new skills by creating directories and writing a new SKILL.md template, which is a broader code-generation/project-modification behavior rather than simply capturing or reviewing learnings.
Using placeholder trigger labels like 'Trigger 1' and 'Trigger 2' gives no model for specificity and normalizes generic invocation rules. This increases the chance that derived skills will be authored with ambiguous triggers, which can cause accidental or excessive skill activation.
No suspicious patterns detected.