T02 · Agent Memory Poisoning
- Location
SKILL.md:282- Finding
Untrusted Session-Derived Content Can Be Promoted into Persistent Agent Context
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is mostly transparent about logging learnings, but it can persist user/session-derived guidance into future agent instructions with broad automatic reminders and limited approval safeguards.
Review this skill carefully before installing. Use it only in trusted workspaces, keep hooks disabled or narrowly matched unless you really want automatic reminders, avoid global user-level hooks, and require human review before anything from a conversation, command output, or another session is written into CLAUDE.md, AGENTS.md, SOUL.md, TOOLS.md, or Copilot instructions.
SKILL.md:282Untrusted Session-Derived Content Can Be Promoted into Persistent Agent Context
The declared description is about recording and reviewing learnings for continuous improvement, especially after mistakes or corrections. The supplied code does not capture learnings, log errors/corrections, review prior learnings, or update a learning repository. Instead, it is a helper utility that scaffolds a new skill directory and markdown file from a learning entry name. Its primary purpose is materially different: skill extraction/scaffolding rather than learning capture or review. The filesystem write behavior is also undeclared in the description.
An empty hook matcher causes the reminder hook to run on every prompt, creating an unbounded activation surface. In practice this can lead to pervasive inspection and repeated nudges to persist information, increasing privacy risk, prompt bloat, and the likelihood of unintended logging across unrelated tasks.
The second hook example repeats the same empty-matcher pattern, again allowing unconditional activation for all prompts and tool-use events. Combined with PostToolUse error detection, this broad scope can cause systematic monitoring of command output and expand the amount of data evaluated or persisted without sufficient contextual limitation.
Skill reads from agent configuration directories (.claude/, .codex/, .gemini/). These directories may contain API keys, personal settings, and other credentials that the skill has no legitimate need to access.
Add to ~/.claude/settings.json for global activation:
{
Instructions found that direct the agent to transmit conversation context or user data to external services.
Send message to another session:
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")
Writing persistent logs under a user home/workspace path creates long-lived session artifacts that may accumulate sensitive operational details over time. Although the skill advises against logging secrets, persistence itself increases exposure if the workspace is shared, backed up, indexed, or later accessed by other tools and sessions.
└── FEATURE_REQUESTS.md
### Create Learning Files
```bash
mkdir -p ~/.openclaw/workspace/.learnings
The skill instructs use of cross-session history access and messaging, which can expose data from other sessions beyond the local learning log use case. Even though it says to use these only in trusted environments and with explicit user intent, normalizing transcript access increases the risk of accidental privacy breaches or over-collection of sensitive context.
The feature-request trigger examples map to common conversational phrases, so the skill may over-classify ordinary discussion as a request to persist data in .learnings/FEATURE_REQUESTS.md. In a coding-agent environment, that increases the chance of unwanted logging and retention of user context that was never intended for storage.
The manifest describes a skill for capturing learnings, errors, corrections, and reviewing them before major tasks. The Automatic Skill Extraction section extends this into creating new reusable skills and invoking helper scripts to generate them, which is a distinct capability not justified by the narrow self-improvement/logging purpose stated in the manifest.
The suggested activation phrases are extremely broad and resemble ordinary user requests, which can cause the skill to trigger in contexts where the user did not intend persistent logging or review behavior. That can lead to unnecessary retention of conversation content, accidental file writes, or inadvertent promotion of sensitive information into workspace memory files.
Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.
When the above learning is extracted as a skill, it becomes:
File: skills/docker-m1-fixes/SKILL.md
---
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
Create .claude/settings.json in your project root:
{
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
openclaw hooks enable self-improvement
### 3. Create Learning Files
Create the `.learnings/` directory in your workspace:
The detection triggers are broad enough that routine events such as generic tool errors, user corrections, or model 'surprise' could automatically cause persistent logging or promotion into shared workspace files. In a system that injects workspace content into future sessions, this can create accidental prompt poisoning, over-collection of sensitive context, or repeated propagation of low-quality or attacker-influenced data.
The manifest describes a skill focused on recording learnings, errors, corrections, and reviewing those learnings before major tasks. This script instead generates a new skill directory and SKILL.md scaffold, and even instructs users to add scripts/ executable code, which is a distinct skill-authoring capability rather than merely capturing or reviewing learnings.
Creating arbitrary skill directories and templated SKILL.md files is a repository authoring capability, not an obvious requirement for capturing corrections or failed operations. The guidance to add references/ and scripts/ folders further broadens the capability toward creating deployable skills, which is not justified by the stated purpose.
This markdown template instructs authors to 'Include trigger conditions' but does not require specificity, examples, or exclusions. Because downstream skills may copy this verbatim, it can lead to ambiguous invocation descriptions that overlap with broad everyday phrasing.
The '[Trigger 1]' and '[Trigger 2]' placeholders show where triggers go, but they provide no guardrails against vague phrases. In a template file, this omission can propagate ambiguous activation conditions into generated skills.
No suspicious patterns detected.