T02 · Agent Memory Poisoning
- Location
SKILL.md:18- Finding
Untrusted Conversational Content Can Be Promoted into Persistent Agent Instructions
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md:18-26, 120-129, 265-289, 446-447
Vulnerability Type: Persistent agent memory poisoning
Risk Level: MediumVulnerable Code Snippets
SKILL.md:18-26:markdown | User corrects you | Log to `.learnings/LEARNINGS.md` with category `correction` | | User wants missing feature | Log to `.learnings/FEATURE_REQUESTS.md` | | API/external tool fails | Log to `.learnings/ERRORS.md` with integration details | | Knowledge was outdated | Log to `.learnings/LEARNINGS.md` with category `knowledge_gap` | | Found better approach | Log to `.learnings/LEARNINGS.md` with category `best_practice` | | Simplify/Harden recurring patterns | Log/update `.learnings/LEARNINGS.md` with `Source: simplify-and-harden` and a stable `Pattern-Key` | | Similar to existing entry | Link with `**See Also**`, consider priority bump | | Broadly applicable learning | Promote to `CLAUDE.md`, `AGENTS.md`, and/or `.github/copilot-instructions.md` | | Workflow improvements | Promote to `AGENTS.md` (OpenClaw workspace) | | Tool gotchas | Promote to `TOOLS.md` (OpenClaw workspace) | | Behavioral patterns | Promote to `SOUL.md` (OpenClaw workspace) |SKILL.md:120-129:markdown ### Add reference to agent files AGENTS.md, CLAUDE.md, or .github/copilot-instructions.md to remind yourself to log learnings. (this is an alternative to hook-based reminders) #### Self-Improvement Workflow When errors or corrections occur: 1. Log to `.learnings/ERRORS.md`, `LEARNINGS.md`, or `FEATURE_REQUESTS.md` 2. Review and promote broadly applicable learnings to: - `CLAUDE.md` - project facts and conventions - `AGENTS.md` - workflows and automation - `.github/copilot-instructions.md` - Copilot contextSKILL.md:265-289:markdown ## Promoting to Project Memory When a learning is broadly applicable (not a one-off fix), promote it to permanent project memory. ### When to Promote - Learning applies across multiple files/features - ...[truncated 3616 chars]- Remediation
View remediation
Remediation Suggestions
- Require explicit workspace-owner approval before modifying any persistent agent-context file.
- Never copy conversational text directly into
CLAUDE.md,AGENTS.md,SOUL.md,TOOLS.md, or Copilot instructions. - Record proposed promotions in a review queue or generated patch rather than applying them automatically.
- Add provenance fields identifying the source session, author, evidence, reviewer, and approval timestamp.
- Verify factual claims against trusted project documentation, code, tests, or authoritative external documentation.
- Treat recurrence only as a prioritization signal, never as evidence that content is trustworthy.
- Reject proposed rules that:
- Override system, developer, or user instructions
- Weaken security checks or approval requirements
- Expand permissions or tool access
- Request secrets or sensitive data
- Introduce external communication or code execution
- Instruct the agent to conceal activity
- Store untrusted learning content as quoted data and clearly delimit it from executable agent instructions.
- Add automated scanning for prompt-injection phrases and security-sensitive directives before promotion.
- Replace “Promote aggressively” with a conservative, review-gated policy requiring verified evidence and least-privilege impact analysis.
