Back to skill

Security audit

M-flow Memory

Security checks for vulnerabilities and agentic risk

Overview

This memory skill appears purpose-aligned rather than malicious, but it needs review because it persistently stores conversations and starts a Docker MCP service with weak controls around network exposure, credentials, and consent.

Install only if you are comfortable with an agent retaining conversation-derived memories across sessions. Prefer running it on a trusted local machine, restrict the Docker port to localhost, use a limited API key with budget controls, review stored memories regularly, and avoid letting it store secrets, regulated data, or confidential work unless you have clear consent and retention controls.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T02 · Agent Memory Poisoning

Error
Location
skill.md:35
Finding

Persistent memory poisoning through automatic conversation storage and retrieval

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/setup.sh:40
Finding

LLM API key exposed during interactive input and retained in Docker metadata

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/setup.sh:76
Finding

Unauthenticated MCP service published on all host network interfaces

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (22)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents this skill as a memory storage/retrieval engine used by agents to remember conversations and preferences. The supplied code chunk does not implement memory storage, retrieval, graph search, or conversational context handling. Instead, it performs cleanup and uninstallation tasks: stopping/removing a Docker container, optionally deleting persisted data, and editing the user's OpenClaw configuration to unregister the service. These are materially different capabilities from the declared purpose and involve destructive system/configuration changes that are not disclosed in the description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents the skill as a memory engine for storing and retrieving long-term conversational context. However, the provided code chunk does not implement memory storage, retrieval, or graph search. Instead, it performs uninstallation/cleanup actions: stopping and removing a Docker container, optionally deleting persisted data, and editing the OpenClaw config to remove the MCP registration. These are materially different capabilities from the declared functional purpose, so this code chunk does not accurately represent the described behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The README explicitly advertises persistent storage of conversations, decisions, preferences, and facts across restarts, but it does not warn users about privacy, retention, or handling of potentially sensitive personal data. In a memory skill whose core function is long-term storage and retrieval of user interactions, omitting these disclosures can lead operators to enable the feature without understanding that sensitive conversation history may be retained indefinitely in Docker volumes and processed by external APIs.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The script collects a sensitive LLM API key from the user and injects it into a third-party container, granting that container access to billable external API usage and any permissions associated with the key. In the context of a memory skill, this is plausible functionality, but it expands trust to the container image and is not clearly bounded or justified with credential-handling safeguards.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

Reading an API key interactively and placing it into the container environment exposes the credential to anyone who can inspect the container configuration, process environment, or related logs on the host. The lack of a clear warning about this exposure increases the chance that users provide a highly privileged key without understanding the risk.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The setup script modifies the user's OpenClaw configuration automatically, altering client behavior and establishing persistence for the new MCP endpoint without explicit confirmation. While this is likely intended for convenience, silent configuration changes can surprise users, overwrite expected state, or register a service they did not fully review.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script updates the user's OpenClaw config without an explicit confirmation prompt, which can silently change how the client connects to MCP services and create hard-to-notice persistence. In a setup script for an agent skill, this is more sensitive because it alters future agent behavior beyond the immediate install step.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
91% confidence
Finding

Maintaining context across sessions is the core feature of the skill, but session persistence inherently increases the blast radius of any mishandled or over-retained data. In this context, cross-session memory is more dangerous because the skill is designed to accumulate and reuse prior conversations, making privacy leakage, stale data reuse, and unintended disclosure more likely.

Content

Scanner excerpt · skill.md (reported line 7)May include surrounding context.

md
Long-term memory engine for OpenClaw agents using M-flow knowledge
  graphs. Stores conversations as structured episodic memories and
  retrieves via graph-routed search. Use when the agent needs to
  remember past conversations, recall user preferences, or maintain
  context across sessions. Requires Docker.
requiredBinaries:
  - docker

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly directs the agent to use persistent long-term memory for conversations and user preferences, but it provides no user-facing consent, retention notice, or data minimization guidance. This creates a privacy risk because sensitive personal information may be stored across sessions and later retrieved or exposed beyond the user's expectations.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The instructions tell the agent to immediately save meaningful conversation content and user-provided facts or preferences whenever the user says to remember something. This is dangerous because it operationalizes automatic retention of natural-language content, which can include sensitive personal, financial, health, or confidential information without classification, confirmation, or sanitization.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The README states that an OpenAI API key is required for LLM-based knowledge extraction, but it does not clearly disclose that conversation-derived content may be sent to an external third-party service for processing. In a memory skill that ingests past conversations and facts, this omission is risky because users may unknowingly transmit sensitive or personal data off-box to an external processor.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The README explicitly advertises persistent storage of conversations, decisions, preferences, and facts across restarts, but it does not warn users about privacy, retention, or sensitivity implications. This can lead operators to enable long-term memory for personal, confidential, or regulated data without informed consent or data-governance controls, increasing the chance of unintended retention and later exposure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The script collects a sensitive LLM API key and injects it into the container as an environment variable. Container environment variables are commonly visible via docker inspect, process metadata, logs, crash reports, or to other local users with Docker access, so this increases the chance of credential exposure.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script silently creates or modifies the user's OpenClaw configuration file to register a new MCP server without prior confirmation. Automatic persistence changes can alter agent behavior and trust boundaries, and if the service later becomes compromised or misconfigured, the agent may connect to an unintended local endpoint on future runs.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
82% confidence
Finding

The skill is explicitly designed to maintain context across sessions by storing past conversations and preferences, which introduces session persistence and cross-session data retention risk. While persistence is the intended feature, it still expands the attack surface for privacy leakage, over-collection, unauthorized recall of prior data, and misuse of historical context if not tightly governed.

Content

Scanner excerpt · skills/mflow-memory/skill.md (reported line 7)May include surrounding context.

md
Long-term memory engine for OpenClaw agents using M-flow knowledge
  graphs. Stores conversations as structured episodic memories and
  retrieves via graph-routed search. Use when the agent needs to
  remember past conversations, recall user preferences, or maintain
  context across sessions. Requires Docker.
---

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill instructs the agent to automatically persist meaningful conversation content and to immediately save user-provided facts when the user says 'remember', but it does not require any explicit privacy notice, consent flow, retention policy, or sensitivity filtering. This creates a real risk of collecting and storing personal, confidential, or regulated data across sessions in a way users may not reasonably expect.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill exposes a destructive prune capability that resets all memory, but the documentation gives no warning about irreversible data loss or any confirmation requirement. In an agent setting, this increases the chance of accidental or unauthorized deletion of long-term memory, which can harm availability and integrity of stored data.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.