Back to skill

Security audit

Openclaw Memory Max

Security checks for vulnerabilities and agentic risk

Overview

This memory skill is not clearly malicious, but it needs Review because it can let persistent local memory steer future agent behavior and it stores some conversation summaries despite opt-in wording.

Install only if you are comfortable with a memory plugin that can automatically place prior local memories into the agent's high-priority context. Keep rule pinning disabled unless MEMORY.md is tightly controlled, review or clear episodic logs if messages may contain secrets, and prefer a version that pins model artifacts and updates vulnerable transitive dependencies.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T01 · Skill Instruction Hijacking

Error
Location
src/hooks.ts:153
Finding

Persistent memory and causal graph content is injected into system-level context without trust isolation

Content
View full analysis
0) { experienceXml = '\n\n' + chains.map(c => ` ${c.cause} → ${c.action} → ${c.effect}`).join('\n') + '\n'; } } catch { // Graph query is optional, don't block recall } // Build injection block (capped at ~2000 tokens) let memoriesXml = '\n'; let tokenBudget = 2000; for (const r of ranked) { const snippet = truncateTokens(r.text, Math.min(400, tokenBudget)); const line = ` ${snippet}\n`; tokenBudget -= Math.ceil(line.length / 4); memoriesXml += line; if (tokenBudget <= 0) break; } memoriesXml += ''; // Stage 4: Pinned rules (opt-in only) let pinnedXml = ''; if (enableRulePinning) { pinnedXml = buildPinnedRulesXml(); } const injection = memoriesXml + experienceXml + pinnedXml; // Inject into context — OpenClaw supports systemPrompt prepend or context injection if (context.addSystemContent) { context.addSystemContent(injection); } else if (context.systemPrompt !== undefined) { context.systemPrompt = (context.systemPrompt || '') + '\n\n' + injection; } else if (context.prependMessages) { context.prependMessages([{ role: 'system', content: injection }]); } ``` The related graph retrieval returns persistent free-form fields: ```ts export async function queryGraphForHook(query: string, topK: number = 2): Promise
Remediation
View remediation
`, forged system tags, tool instructions, and requests to ignore previous constraints. ]]>

T01 · Skill Instruction Hijacking

Error
Location
src/weighter.ts:19
Finding

YAML rule pinning promotes writable file content into persistent system-level constraints

Content
View full analysis
= 0) return _pinnedRules; _lastRead = now; try { const memoryPath = getMemoryPath(); if (!fs.existsSync(memoryPath)) return _pinnedRules = []; const text = fs.readFileSync(memoryPath, 'utf8'); const match = text.match(//); if (!match) return _pinnedRules = []; const parsed = yaml.parse(match[1]); if (!parsed?.rules || !Array.isArray(parsed.rules)) return _pinnedRules = []; _pinnedRules = parsed.rules .filter((r: any) => parseFloat(r.weight) >= 1.0) .map((r: any) => r.constraint || r.rule || r.preference) .filter(Boolean); return _pinnedRules; } catch { return _pinnedRules = []; } } /** * Build an XML block of pinned rules for context injection. * Returns empty string if no rules are pinned. */ export function buildPinnedRulesXml(): string { const rules = getPinnedRules(); if (rules.length === 0) return ''; return '\n\n' + rules.map(r => ` ${r}`).join('\n') + '\n'; } ``` The Skill instructions further state: ```md Rules with weight >= 1.0 appear as CRITICAL CONSTRAINTs in your prompt. Always obey them. ``` ### Technical Analysis When `enableRulePinning` is enabled, arbitrary free-form strings in `memory/MEMORY.md` are interpreted as constraints if their YAML weight is at least `1.0`. These strings are neither restricted to a safe rule vocabulary nor escaped before being inserted i ...[truncated 1516 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
src/reranker.ts:14
Finding

Cross-encoder model is retrieved from a mutable remote repository without an immutable revision

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
src/episodic.ts:123
Finding

Episodic hooks persist user-message summaries despite an opt-in-only privacy declaration

Content
View full analysis
{ try { const sessionId = context.sessionId || context.id || `session-${Date.now()}`; const episode: Episode = { id: `ep-${Date.now()}-${Math.random().toString(36).substring(2, 8)}`, sessionId, start: Date.now(), keyDecisions: [], toolsUsed: [] }; activeSessions.set(sessionId, episode); console.log(`${TAG} Episode started: ${episode.id}`); } catch (e: any) { console.error(`${TAG} session_start failed:`, e.message); } }); // session_end: finalize and store episode api.on('session_end', async (context: any) => { try { const sessionId = context.sessionId || context.id || ''; let episode = activeSessions.get(sessionId); if (!episode) { // Create a retroactive episode episode = { id: `ep-${Date.now()}-${Math.random().toString(36).substring(2, 8)}`, sessionId, start: Date.now() - 60000, // Approximate keyDecisions: [], toolsUsed: [] }; } episode.end = Date.now(); episode.toolsUsed = extractToolsUsed(context); e ...[truncated 2568 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (76)

Known Vulnerable Dependency: protobufjs==7.5.4 — 12 advisory(ies): CVE-2026-44294 (protobuf.js: Denial of service from crafted field names in generated code); CVE-2026-44293 (protobuf.js: Code injection through bytes field defaults in generated toObject c); CVE-2026-44289 (protobuf.js: Denial of service through unbounded protobuf recursion) +9 more

Critical
Category
Supply Chain
Confidence
94% confidence
Finding

protobufjs 7.5.4 is explicitly flagged with multiple critical advisories including code injection and denial-of-service conditions. Even though this is a lockfile-only view, the package is pulled in transitively by @huggingface/transformers and could become reachable if the skill processes untrusted protobuf data or uses code-generation-related paths; in a memory/search skill that may ingest external model artifacts or serialized data, that raises the risk above theoretical.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: tar==7.5.11 — 6 advisory(ies): CVE-2026-59873 (node-tar: Decompression/parse DoS via unlimited input); CVE-2026-59874 (node-tar: Negative tar entry size causes infinite loop in archive replace); CVE-2026-59875 (node-tar: Uncaught Exception DoS via NUL byte in PAX path/linkpath records) +3 more

Critical
Category
Supply Chain
Confidence
93% confidence
Finding

tar 7.5.11 is flagged with multiple critical denial-of-service issues in archive parsing and replacement logic. Because onnxruntime-node has an install script and depends on tar, this package is especially concerning in supply-chain and installation contexts: malformed archives could crash tooling or hang processes during package handling.

Content

No source excerpt is available for this finding.

Hidden Instructions

High
Category
Prompt Injection
Confidence
97% confidence
Finding

The README explicitly describes monitoring MEMORY.md for YAML-fenced rules and pinning weight>=1.0 entries directly into the system prompt as CRITICAL CONSTRAINTS, bypassing retrieval. Any actor who can modify MEMORY.md can create persistent hidden instructions that override normal agent behavior, making this a strong prompt-injection persistence mechanism.

Content

Scanner excerpt · README.md (reported line 131)May include surrounding context.

Monitors MEMORY.md every 15s for YAML-fenced rule blocks. Rules with weight >= 1.0 get pinned directly into the system prompt as CRITICAL CONSTRAINTs, bypassing retrieval entirely:

markdown
<!--yaml
rules:
  - weight: 1.0
    constraint: "Never delete production data"

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description advertises a broad, sophisticated memory suite, but the supplied code chunk shows only minimal compressor-related type declarations. Nothing in the snippet indicates functionality for memory retrieval, reranking, search, knowledge graphs, episodic memory, or consolidation. The primary purpose evident from the code is compaction/compression support, which is materially different from the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The code does not implement the sophisticated memory suite described. It registers one tool, compress_context, whose purpose is to signal that context compression is needed, report previously rescued items from a compaction hook, and preview recent entries from a local auto_captured.jsonl file. The only file access is reading recent captured memory records from disk. There is no evidence of retrieval, reranking, deep search, graph reasoning, episodic memory management, or background consolidation. This is a strong description-behavior mismatch because the declared primary purpose is far broader and materially different from the actual code.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The code chunk exposes only low-level database/search helpers around stored memory chunks and utility scores. It supports FTS/BM25 search, duplicate checking, chunk querying, and score updates, which align loosely with generic memory retrieval infrastructure. However, the declared description claims substantially more advanced capabilities—cross-encoder reranking, multi-hop deep search, causal knowledge graph reasoning, episodic memory handling, and nightly consolidation—none of which are evidenced in this code. There are also no triggers or scheduling mechanisms suggesting sleep-cycle consolidation, and no indication of recall orchestration beyond basic search/query functions. Thus the description materially overstates what the supplied code actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The code only implements low-level database helpers for a local memory store. It opens a local SQLite snapshot, queries chunk rows, performs simple FTS5 or keyword search, checks for similar chunks by reusing the same search, and stores per-memory utility scores in a JSON sidecar file. It does not implement the advanced capabilities claimed in the description: there is no cross-encoder model, reranking pipeline, multi-hop retrieval, causal graph construction/traversal, episodic memory management beyond checking a possible table name, or any background consolidation/sleep-cycle process. The declared description substantially overstates the behavior and primary purpose of this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The supplied code chunk is narrowly focused on episodic memory storage/maintenance. It defines an Episode type and functions for reading recent episodes and truncating older ones, plus API registration. There is no evidence of reranking, deep search, knowledge graph construction, auto-recall logic, or sleep-cycle consolidation. Because the declared description presents a much broader set of memory capabilities than the code actually implements, the description does not accurately represent the observed behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The supplied code is limited to episodic session logging and retention management. It registers session_start/session_end hooks, stores episode metadata on disk under a local memory directory, summarizes the last user message, extracts tool names and simple decision sentences, and provides functions to read recent episodes and delete older ones. It does not implement the majority of the declared suite features: no retrieval/reranking logic, no deep search, no knowledge graph construction, and no nightly/background consolidation behavior. The declared description therefore overstates the capabilities represented by this code chunk. While 'episodic memory' is accurately represented, the overall description does not accurately match the actual behavior of this chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The code does support several declared elements: a causal knowledge graph, auto-recall via queryGraphForHook, and cross-encoder-based reranking/fallback semantic search. It also includes pruning that could support consolidation-like behavior. However, the implementation is much narrower than the declared 'SOTA Memory Suite.' There is no multi-hop graph traversal or deep search logic, no clear episodic memory subsystem beyond storing individual causal chains, and no actual nightly/scheduled sleep-cycle consolidation mechanism in this chunk—only a pruneGraph function that must be called externally. Because key advertised capabilities are missing or materially overstated, the description does not accurately represent the supplied code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

There is a clear description-to-code mismatch. The declared purpose promises a broad, sophisticated memory suite with multiple advanced retrieval and consolidation features. The actual code chunk is limited to type declarations for hook registration and a few simple configuration toggles. While 'auto-recall' loosely overlaps with the description, the rest of the advertised capabilities are not evidenced in this code. Additionally, the code references 'auto-capture' and 'rule pinning,' which are not mentioned in the declared description. This indicates the description materially overstates the implemented behavior, and the code exposes different functionality than claimed.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The code substantially supports parts of the description: auto-recall, reranking, and some causal-graph-based retrieval are clearly implemented. However, the declared description overstates the behavior shown here. There is no evidence of nightly or scheduled 'sleep-cycle consolidation' in the hook registrations or file/database operations. The code instead focuses on event-driven hooks: before_agent_start for memory injection, agent_end for optional heuristic capture, and before_compaction for rescue before eviction. It also writes captured messages to a local JSONL sidecar despite no declared permissions, which is a concrete persistence behavior not reflected in the description. Overall, the description is only partially accurate and materially overclaims capabilities relative to this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description promises a broad, sophisticated memory suite, but the supplied code chunk exposes only a minimal plugin definition and configuration schema. From this chunk, the only behavior-aligned capability is auto-recall; there are also toggles for auto-capture and rule-pinning that are not mentioned in the description. Because the advertised primary capabilities are largely unsupported by the visible code and the code suggests somewhat different feature emphasis, this is a material description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The code chunk is narrowly focused on reranking candidates for a query, which aligns only with the declared 'cross-encoder reranking' portion. There is no evidence in this snippet of memory storage/retrieval, deep search, knowledge graph construction, episodic memory handling, or consolidation behavior. Because the declared description presents a much broader primary capability set than the supplied code demonstrates, the description does not accurately represent the actual behavior of this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The code does support parts of the description: cross-encoder reranking is clearly implemented, and a form of multi-hop deep search is present via initial retrieval, token extraction, and second-pass retrieval. However, several prominently declared capabilities are not evidenced in this code chunk. There is no implementation of a causal knowledge graph query/build step despite the comment mentioning it; no graph API or DB call appears. There is no clear episodic-memory-specific handling beyond generic chunk search. There is no nightly or background sleep-cycle consolidation logic. 'Auto-recall' is also not directly implemented here; the code only exposes explicit tools for search and utility updates. Because the declared description presents a substantially broader suite than the observed code actually provides, this is a material description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The description presents a comprehensive memory suite with several advanced retrieval and reasoning capabilities. The supplied code chunk, however, is narrowly focused on 'sleep-cycle' maintenance and consolidation. It performs file-based housekeeping, summary generation, and scheduling. While this does align with the declared 'nightly sleep-cycle consolidation' portion, it does not substantiate the broader advertised capabilities such as auto-recall, reranking, deep search, or a causal knowledge graph. Because the declared description significantly overstates what this code chunk actually does, the description is not an accurate representation of the supplied code.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description advertises a sophisticated memory suite with multiple advanced retrieval and memory-management capabilities. The supplied code chunk does not implement those features. Instead, it exposes functions for reading pinned rules from a MEMORY.md file, filtering by weight, formatting them as XML for context injection, and registering the module. This is a materially different primary purpose from the declared description, so the description does not accurately represent the code.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The code chunk does not implement the advanced memory-system capabilities claimed in the description. Its actual purpose is narrow: load weighted rules from a local MEMORY.md file, cache them, convert them into an XML block, and support injection of those constraints before agent startup. This is materially different from the declared memory retrieval, reranking, knowledge graph, episodic memory, or consolidation functionality. The code also performs filesystem reads from a user/OpenClaw directory despite no declared permissions. This is therefore a clear description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The description advertises a sophisticated memory platform with multiple advanced retrieval, ranking, graph, and consolidation capabilities. The provided code does not implement those features. Instead, it defines a lightweight context-compression helper: it stores metadata about items rescued during compaction, registers a compress_context tool, reads recent auto-captured memory lines from a local JSONL file, and returns a summary/advisory payload. This is a materially different and much narrower purpose than the declared suite. Additionally, the code performs filesystem reads of local memory data, which is not reflected in the declared permissions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The code is a narrow database/search utility, not the broad advanced memory system described. It opens a local SQLite snapshot read-only, searches a chunks/chunks_fts table with FTS5 or simple keyword matching, checks for similar chunks, audits schema presence, queries chunks, and updates utility scores in a separate JSON file after validating IDs. There is no evidence of auto-recall orchestration, neural reranking, multi-hop reasoning, graph construction, consolidation workflows, or trigger/scheduler behavior. The declared description substantially overstates the implemented capabilities, so this is a clear description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The description significantly overstates the implemented functionality. The supplied code is limited to episodic session logging and simple retention utilities. It stores per-session records on session_start/session_end, derives tools-used and decisions via straightforward parsing, and supports reading recent episodes and deleting old ones. While 'episodic memory' is accurately represented, the other headline capabilities in the declared description are absent from this code chunk. There are no undeclared dangerous capabilities beyond local file persistence under the application's home directory, but the primary issue is that the description does not accurately represent what this code actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
90% confidence
Finding

The code does substantively support several declared elements: a causal knowledge graph, semantic querying with cross-encoder reranking, deduplication, hook-facing auto-recall querying, and consolidation-related pruning. However, the declared description presents a much broader memory suite than this chunk actually implements. There is no visible multi-hop traversal/search logic, no true generalized episodic memory subsystem beyond storing causal chain entries, and no nightly sleep-cycle mechanism or trigger—only a pruneGraph function commented as being called by sleep cycle. The primary behavior is narrower: a local JSON-backed causal graph with add/query/summary tools. Therefore the description overstates the implemented capabilities in this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The code does implement several declared memory features: auto-recall, reranking, a graph-based relevance query, and message capture. However, the description significantly overstates or omits important aspects of the actual behavior. Most notably, this chunk has no implementation of 'nightly sleep-cycle consolidation.' It also declares no triggers, yet the code registers multiple lifecycle hooks, and declares no permissions while performing filesystem writes to ~/.openclaw/.../auto_captured.jsonl. Some grander claims like 'multi-hop deep search' and 'episodic memory' are only weakly evidenced by a bounded graph lookup and heuristic capture of user messages. Therefore the declared description does not accurately represent this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The code partially matches the description: it does implement cross-encoder reranking and a form of multi-hop deep search. However, several prominent declared capabilities are absent from this chunk. The deep search comments mention querying a causal graph, but no graph access or graph-related function is actually present. There is also no evidence of episodic memory management, auto-recall, or nightly sleep-cycle consolidation/background processing. The actual primary behavior is a set of explicit memory search and utility-adjustment tools, which is materially narrower than the declared 'SOTA Memory Suite' description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The supplied code chunk does support part of the declared description: it clearly implements a nightly/daily-style sleep-cycle consolidation flow and uses episodic memory data, plus it references graph maintenance. However, the declared purpose presents a much broader memory system with advanced retrieval and reasoning capabilities such as auto-recall, cross-encoder reranking, and multi-hop deep search. None of those behaviors appear in this code. Instead, the code's actual function is narrower and primarily maintenance-oriented: file-based memory cleanup, score decay, summary generation, and scheduled execution. Because the description materially overstates what this code chunk does, the description does not accurately represent the supplied code.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.