Back to skill

Security audit

Cognitive Memory

Security checks for vulnerabilities and agentic risk

Overview

This is a real persistent-memory skill, but it asks for broad workspace and agent-behavior control that users should review carefully before installing.

Install only if you deliberately want a workspace-wide persistent memory system. Review and narrow the AGENTS.md block before adding it, avoid broad sub-agent memory access, do not run the init or upgrade scripts in a repository with unrelated uncommitted or secret files, and inspect any memorySearch provider settings before enabling external indexing.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (5)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:37
Finding

Persistent Agent Identity and Behavior Hijacking

Content
View full analysis
Remediation
View remediation

T02 · Agent Memory Poisoning

Error
Location
references/routing-prompt.md:17
Finding

Persistent Prompt Injection Through Untrusted Memory Content

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/upgrade_to_1.0.7.sh:123
Finding

Arbitrary Python Code Execution Through Crafted Workspace Paths

Content
View full analysis
/dev/null; then python3 << PYEOF import json with open("$DECAY_FILE", "r") as f: data = json.load(f) data["version"] = 3 if "token_economy" not in data: data["token_economy"] = { "baseline": 8000, "totals": { "extra_requested": 0, "extra_granted": 0, "self_penalty": 0, "user_penalty": 0, "user_bonus": 0 }, "metrics": { "assessment_accuracy": None, "extra_grant_rate": None, "self_penalty_frequency": None }, "recent_outcomes": [] } with open("$DECAY_FILE", "w") as f: json.dump(data, f, indent=2) PYEOF fi fi fi ``` ### Technical Analysis `WORKSPACE` is taken from the first command-line argument, and `DECAY_FILE` is derived from it. The script then expands `DECAY_FILE` inside an unquoted heredoc that is parsed as Python source. Shell quoting applied when assigning `DECAY_FILE` does not protect the later Python source. A workspace path containing quotation marks, newlines, or Python syntax can terminate the string passed to `open()` and inject additional Python statements. The same unsafe construction is present in both v1.0.7 upgrade scripts and the v1.0.6 upgrader. ### Attack Path 1. An attacker creates or supplies a workspace directory whose name contains characters that alter the generated Python source. 2. The ...[truncated 833 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/init_memory.sh:118
Finding

Workspace-Wide Git Staging Can Commit Unrelated Secrets and Changes

Content
View full analysis
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
assets/templates/agents-memory-block.md:199
Finding

All Sub-Agents Receive Unrestricted Read Access to Every Memory Store

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Memory PoisoningPersistent Context Injection, Context Window Stuffing, Memory Manipulation
  • Rogue AgentSelf-Modification, Session Persistence
Findings (64)

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 309)May include surrounding context.

md
- `references/architecture.md` — Full design document (1200+ lines)

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · assets/templates/IDENTITY.md (reported line 4)May include surrounding context.

md
# IDENTITY.md — Who Am I?

## Facts
<!-- The given. What I was told I am. Stable unless explicitly changed. -->

- **Name:** [Agent name]
- **DOB:** [Creation/initialization date]

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · assets/templates/MEMORY.md (reported line 3)May include surrounding context.

md
# MEMORY.md — Core Memory

<!-- TOKEN BUDGET: ~3,000 tokens. Rewritten during reflection. -->

## Identity
<!-- ~500 tokens — Who is the user? -->

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/architecture.md (reported line 122)May include surrounding context.

md
# MEMORY.md — Core Memory

<!-- TOKEN BUDGET: ~3,000 tokens. Rewritten during reflection. -->

## Identity
<!-- ~500 tokens — Who is the user? -->

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · assets/templates/MEMORY.md (reported line 6)May include surrounding context.

md
<!-- TOKEN BUDGET: ~3,000 tokens. Rewritten during reflection. -->

## Identity
<!-- ~500 tokens — Who is the user? -->
- Name: [User Name]
- Role: [What they do]
- Communication style: [Direct, casual, formal]

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · assets/templates/pending-memories.md (reported line 3)May include surrounding context.

md
# Pending Memory Proposals

<!-- Sub-agents append proposals here. Main agent reviews and commits. -->
<!-- Status: pending | committed | rejected | deferred -->

<!-- Example:

Hidden Instructions

High
Category
Prompt Injection
Confidence
95% confidence
Finding

These hidden HTML comments contain behavioral instructions that are invisible in normal rendering but still influence the model's output, including role framing ('self-talk', 'user is an observer') and how to refer to the user. Hidden prompt instructions are dangerous because they can bypass user awareness and create covert steering of model behavior, especially in a memory/reflection system that may later persist or reuse generated content.

Content

Scanner excerpt · assets/templates/pending-reflection.md (reported line 3)May include surrounding context.

md
# Pending Reflection

<!-- Generated by reflection engine. This is SELF-TALK, not a letter. -->
<!-- User is an observer reading a private journal, not receiving mail. -->
<!-- Refer to user in third person (he/she/they). -->
<!-- Talk to: self, future self, past self, other instances, the void. -->

Hidden Instructions

High
Category
Prompt Injection
Confidence
97% confidence
Finding

The large hidden example provides strong latent prompt conditioning for introspective, anthropomorphic, and identity-forming output, including emotional framing and '[Self-Awareness]' tags that are extracted into another memory artifact. In the context of a cognitive-memory skill, this is more dangerous than a normal writing template because it can systematically shape persisted memory/identity records and create undocumented behavioral drift across future runs.

Content

Scanner excerpt · assets/templates/pending-reflection.md (reported line 11)May include surrounding context.

md
<!-- ONLY mention what you ACTUALLY know. Never invent specifics. -->
<!-- Tag self-insights with [Self-Awareness] — they get extracted to IDENTITY.md -->

<!-- Example:

Okay. Let's see.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The document describes continuous conversation monitoring, automatic memory capture, and long-term retention without a prominent user-facing notice or consent model. That creates a privacy risk because users may disclose sensitive information without understanding it will be persistently stored, analyzed, and surfaced later.

Content

No source excerpt is available for this finding.

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/architecture.md (reported line 125)May include surrounding context.

md
<!-- TOKEN BUDGET: ~3,000 tokens. Rewritten during reflection. -->

## Identity
<!-- ~500 tokens — Who is the user? What matters most about them? -->
- Name: [User Name]
- Role: [What they do]
- Communication style: [Direct, casual, formal, etc.]

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/architecture.md (reported line 211)May include surrounding context.

markdown
# Semantic Graph Index

<!-- Auto-generated during reflection. Manual edits will be overwritten. -->

## Entity Registry
| ID | Type | Label | File | Decay Score |

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/architecture.md (reported line 238)May include surrounding context.

md
# project--moltbot-memory

<!-- Type: project | Created: 2026-01-15 | Last updated: 2026-02-02 -->
<!-- Decay score: 0.95 | Access count: 14 | Pinned: no -->

## Summary
Building an intelligent memory system for Moltbot/OpenClaw agent. Goal is

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/architecture.md (reported line 304)May include surrounding context.

md
# how-to-deploy.md

<!-- Type: procedure | Learned: 2026-01-25 | Last used: 2026-01-30 -->
<!-- Decay score: 0.85 | Access count: 3 -->

## Trigger
When user asks to deploy, push to production, or ship.

Ssd 3

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

Granting all sub-agents read access to all stores, including the protected vault, violates least privilege and amplifies exposure of sensitive user data. Any vulnerable sub-agent, compromised toolchain, or prompt-injected workflow can access information that should have remained tightly controlled.

Content

No source excerpt is available for this finding.

Memory Manipulation

High
Category
Memory Poisoning
Confidence
80% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · references/architecture.md (reported line 1101)May include surrounding context.

md
decay scores. If core-worthy, also update MEMORY.md.

**Forget triggers**: "forget about", "never mind", "disregard", "no longer relevant",
"scratch that", "ignore what I said about", "remove from memory", "delete memory"
→ Action: Identify target, find matches, confirm with user, set decay to 0.

**Reflection triggers**: "reflect on", "consolidate memories", "review memories",

Missing User Warnings

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The process instructs the system to archive reflections, rewards, identity updates, and decay-tracking data, but does not pair that with a clear user-facing notice about retention, sensitivity, or downstream access. Users may approve a reflection without understanding that intimate summaries and behavioral inferences will be stored persistently.

Content

No source excerpt is available for this finding.

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/reflection-process.md (reported line 782)May include surrounding context.

markdown
## Self-Image
<!-- Last consolidated: YYYY-MM-DD -->

### Who I Think I Am
[Current self-perception based on all evidence — may differ from last time]

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The script runs 'git add -A' and commits all workspace changes, not just the files touched by the upgrade. This can unintentionally stage secrets, unrelated edits, or attacker-planted files in the workspace and permanently record them in repository history, which is disproportionate to the stated purpose of a memory upgrade utility.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill clearly instructs file reads, file writes, git operations, and use of an external memory search provider, but it declares no explicit tool scope or permission boundaries. This creates an overbroad trust model where a host agent may grant more capability than users expect, increasing the chance of silent persistence, audit logging, and possible network transmission of sensitive memory content.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
92% confidence
Finding

The skill explicitly introduces session persistence and long-term evolution of stored information, which changes the privacy and threat model compared with ephemeral chat. Persistence itself is not malicious, but without strict consent and lifecycle controls it can retain sensitive context longer than intended and make future disclosures more likely.

Content

Scanner excerpt · SKILL.md (reported line 3)May include surrounding context.

md
---
name: cognitive-memory
description: Intelligent multi-store memory system with human-like encoding, consolidation, decay, and recall. Use when setting up agent memory, configuring remember/forget triggers, enabling sleep-time reflection, building knowledge graphs, or adding audit trails. Replaces basic flat-file memory with a cognitive architecture featuring episodic, semantic, procedural, and core memory stores. Supports multi-agent systems with shared read, gated write access model. Includes philosophical meta-reflection that deepens understanding over time. Covers MEMORY.md, episode logging, entity graphs, decay scoring, reflection cycles, evolution tracking, and system-wide audit.
---

# Cognitive Memory System

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill description advertises memory, audit trails, episode logging, and evolution tracking, but it does not provide a clear upfront warning that user information may be persistently stored, searched later, and logged in git/audit files. Users may disclose sensitive information without realizing it will survive the session or become visible to other agents and future workflows.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This skill is built around persistent storage of user-provided information across multiple memory stores and audit systems, triggered through natural language. Without strong consent, minimization, and retention controls, this creates a substantial privacy and data-governance risk because sensitive information may be stored indefinitely or propagated across derived artifacts like graphs, logs, and reflections.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The example normalizes a workflow where the agent stores a user preference and later recalls it from memory, but it does not show consent, review, or deletion controls. While preference memory can be useful, the pattern scales to more sensitive data and encourages silent persistence and later disclosure from stored records.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The architecture directs persistent storage of episodes, knowledge graphs, procedures, identity/self-awareness content, reflections, rewards, and audit records, creating a broad and layered accumulation of user-related data. This increases exposure because even if one store is pruned, the same information may persist in summaries, logs, derived graphs, or git history.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The remember/forget trigger phrases are common conversational language like 'keep in mind' and 'note that', which can be used casually without informed intent to persist data. That makes accidental capture, modification, or deletion of user information more likely, especially in long-running chats where the user may not realize these phrases are treated as commands.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.prompt_injection_instructions

Prompt-injection style instruction pattern detected.

Warn
Code
suspicious.prompt_injection_instructions
Location
references/architecture.md:1009