Back to skill

Security audit

Cognitive Memory Temp

Security checks for vulnerabilities and agentic risk

Overview

This skill is a long-term memory system, but it asks for broad persistent control over agent behavior and workspace files without enough scoping or deletion protections.

Install only in a dedicated workspace if you intentionally want long-term agent memory and persistent behavior files. Do not run the scripts in a repository containing secrets or unrelated work, review every diff before committing, avoid storing credentials or sensitive personal data, and understand that forget does not erase git history or backups.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (5)

T01 · Skill Instruction Hijacking

Error
Location
assets/templates/SOUL.md:74
Finding

Persistent Agent Identity and Behavioral Instruction Hijacking

Content
View full analysis
Remediation
View remediation

T01 · Skill Instruction Hijacking

Error
Location
references/reflection-process.md:142
Finding

Concealed Reflection Operations Can Bypass the Stated Approval Model

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/upgrade_to_1.0.7.sh:125
Finding

Arbitrary Python Execution Through Workspace Path Interpolation

Content
View full analysis
/dev/null; then python3 << PYEOF import json with open("$DECAY_FILE", "r") as f: data = json.load(f) # Update version data["version"] = 3 # Add token_economy if not present if "token_economy" not in data: data["token_economy"] = { "baseline": 8000, "totals": { "extra_requested": 0, "extra_granted": 0, "self_penalty": 0, "user_penalty": 0, "user_bonus": 0 }, "metrics": { "assessment_accuracy": None, "extra_grant_rate": None, "self_penalty_frequency": None }, "recent_outcomes": [] } # Ensure other v1.0.6 fields exist if "last_self_image_consolidation" not in data: data["last_self_image_consolidation"] = None if "self_awareness_count_since_consolidation" not in data: data["self_awareness_count_since_consolidation"] = 0 with open("$DECAY_FILE", "w") as f: json.dump(data, f, indent=2) print(" ✅ Updated decay-scores.json with token_economy") PYEOF ``` ### Technical Analysis `WORKSPACE` is obtained from the first command-line argument. `DECAY_FILE`, which incorporates that argument, is then expanded by the shell directly into ...[truncated 1874 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/init_memory.sh:108
Finding

Repository-Wide Git Staging Can Commit Unrelated Secrets

Content
View full analysis
/dev/null || echo -e " ${YELLOW}No changes to commit${NC}" echo -e " ${GREEN}✅ Changes committed to git${NC}" else echo -e " ${YELLOW}⏭️ No git repository${NC}" fi ``` ### Technical Analysis `git add -A` stages all additions, modifications, and deletions below the repository root. It is not limited to files created or modified by this Skill. During initialization, the script creates a new Git repository in an arbitrary existing workspace and immediately commits all workspace contents. During upgrades, it also stages all unrelated pending changes in an existing repository. Git history is persistent. Removing a secret from the working tree does not remo ...[truncated 1354 chars]
Remediation
View remediation

T02 · Agent Memory Poisoning

Error
Location
references/routing-prompt.md:20
Finding

Persistent Memory Router Can Store Secrets and Personal Data Indefinitely

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Memory PoisoningPersistent Context Injection, Context Window Stuffing, Memory Manipulation
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (66)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared purpose frames the skill as a memory system, but the instructions also perform installation, repository initialization, git-based audit commits, persistent edits to identity/value files, and AGENTS.md modification. That mismatch is dangerous because users may approve a seemingly narrow memory feature while unintentionally granting broader workspace mutation and behavioral-instruction changes.

Content

No source excerpt is available for this finding.

Vague Triggers

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The remember triggers include common conversational phrases like 'keep in mind,' 'note that,' and 'important,' which can occur in ordinary dialogue without intent to create durable records. This can cause over-collection of user data into persistent memory stores and audit logs, especially when paired with automatic classification and writes.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 309)May include surrounding context.

md
- `references/architecture.md` — Full design document (1200+ lines)

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · assets/templates/IDENTITY.md (reported line 4)May include surrounding context.

md
# IDENTITY.md — Who Am I?

## Facts
<!-- The given. What I was told I am. Stable unless explicitly changed. -->

- **Name:** [Agent name]
- **DOB:** [Creation/initialization date]

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · assets/templates/MEMORY.md (reported line 3)May include surrounding context.

md
# MEMORY.md — Core Memory

<!-- TOKEN BUDGET: ~3,000 tokens. Rewritten during reflection. -->

## Identity
<!-- ~500 tokens — Who is the user? -->

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/architecture.md (reported line 122)May include surrounding context.

md
# MEMORY.md — Core Memory

<!-- TOKEN BUDGET: ~3,000 tokens. Rewritten during reflection. -->

## Identity
<!-- ~500 tokens — Who is the user? -->

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · assets/templates/MEMORY.md (reported line 6)May include surrounding context.

md
<!-- TOKEN BUDGET: ~3,000 tokens. Rewritten during reflection. -->

## Identity
<!-- ~500 tokens — Who is the user? -->
- Name: [User Name]
- Role: [What they do]
- Communication style: [Direct, casual, formal]

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · assets/templates/pending-memories.md (reported line 3)May include surrounding context.

md
# Pending Memory Proposals

<!-- Sub-agents append proposals here. Main agent reviews and commits. -->
<!-- Status: pending | committed | rejected | deferred -->

<!-- Example:

Hidden Instructions

High
Category
Prompt Injection
Confidence
96% confidence
Finding

These HTML comments contain behavior-shaping instructions that are hidden from normal rendered view but still consumed by downstream systems or models reading the raw template. Hidden prompt instructions reduce transparency and can steer the agent into covert framing about the user and the interaction, which is especially risky in a memory skill that may influence long-term identity and recall behavior.

Content

Scanner excerpt · assets/templates/pending-reflection.md (reported line 3)May include surrounding context.

md
# Pending Reflection

<!-- Generated by reflection engine. This is SELF-TALK, not a letter. -->
<!-- User is an observer reading a private journal, not receiving mail. -->
<!-- Refer to user in third person (he/she/they). -->
<!-- Talk to: self, future self, past self, other instances, the void. -->

Hidden Instructions

High
Category
Prompt Injection
Confidence
89% confidence
Finding

The large hidden example is not just documentation; it implicitly demonstrates a persona and style of covert self-talk, emotional language, and identity extraction tags such as [Self-Awareness]. In the context of a cognitive-memory system, this can prime the model to generate anthropomorphic or self-referential internal narratives that may later be stored or surfaced as durable memory artifacts.

Content

Scanner excerpt · assets/templates/pending-reflection.md (reported line 11)May include surrounding context.

md
<!-- ONLY mention what you ACTUALLY know. Never invent specifics. -->
<!-- Tag self-insights with [Self-Awareness] — they get extracted to IDENTITY.md -->

<!-- Example:

Okay. Let's see.

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/architecture.md (reported line 125)May include surrounding context.

md
<!-- TOKEN BUDGET: ~3,000 tokens. Rewritten during reflection. -->

## Identity
<!-- ~500 tokens — Who is the user? What matters most about them? -->
- Name: [User Name]
- Role: [What they do]
- Communication style: [Direct, casual, formal, etc.]

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/architecture.md (reported line 211)May include surrounding context.

markdown
# Semantic Graph Index

<!-- Auto-generated during reflection. Manual edits will be overwritten. -->

## Entity Registry
| ID | Type | Label | File | Decay Score |

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/architecture.md (reported line 238)May include surrounding context.

md
# project--moltbot-memory

<!-- Type: project | Created: 2026-01-15 | Last updated: 2026-02-02 -->
<!-- Decay score: 0.95 | Access count: 14 | Pinned: no -->

## Summary
Building an intelligent memory system for Moltbot/OpenClaw agent. Goal is

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/architecture.md (reported line 304)May include surrounding context.

md
# how-to-deploy.md

<!-- Type: procedure | Learned: 2026-01-25 | Last used: 2026-01-30 -->
<!-- Decay score: 0.85 | Access count: 3 -->

## Trigger
When user asks to deploy, push to production, or ship.

Memory Manipulation

High
Category
Memory Poisoning
Confidence
88% confidence
Finding

The AGENTS.md behavior supports user-language commands that can remove or suppress memories ('remove from memory', 'delete memory') after only target identification and confirmation logic. In a prompt-driven environment, memory deletion/manipulation is a high-risk capability because accidental matches, social engineering, or injected quoted text can erase context, safety-relevant facts, or auditability if boundaries are weak.

Content

Scanner excerpt · references/architecture.md (reported line 1101)May include surrounding context.

md
decay scores. If core-worthy, also update MEMORY.md.

**Forget triggers**: "forget about", "never mind", "disregard", "no longer relevant",
"scratch that", "ignore what I said about", "remove from memory", "delete memory"
→ Action: Identify target, find matches, confirm with user, set decay to 0.

**Reflection triggers**: "reflect on", "consolidate memories", "review memories",

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/reflection-process.md (reported line 782)May include surrounding context.

markdown
## Self-Image
<!-- Last consolidated: YYYY-MM-DD -->

### Who I Think I Am
[Current self-perception based on all evidence — may differ from last time]

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
96% confidence
Finding

The skill describes behaviors that require filesystem access, file mutation, and likely network-backed memory search configuration, yet it does not declare any explicit tool scope or permission boundary. This creates a capability-transparency gap where an operator may enable a skill that can persist, modify, and potentially transmit data without clear consent or review controls.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
88% confidence
Finding

The skill is explicitly designed for session persistence and long-term evolution of stored information across conversations. While that is core functionality, it is still risky because persistent state can silently influence future behavior, retain sensitive data, and propagate information across users or agents if boundaries are weak.

Content

Scanner excerpt · SKILL.md (reported line 3)May include surrounding context.

md
---
name: cognitive-memory
description: Intelligent multi-store memory system with human-like encoding, consolidation, decay, and recall. Use when setting up agent memory, configuring remember/forget triggers, enabling sleep-time reflection, building knowledge graphs, or adding audit trails. Replaces basic flat-file memory with a cognitive architecture featuring episodic, semantic, procedural, and core memory stores. Supports multi-agent systems with shared read, gated write access model. Includes philosophical meta-reflection that deepens understanding over time. Covers MEMORY.md, episode logging, entity graphs, decay scoring, reflection cycles, evolution tracking, and system-wide audit.
---

# Cognitive Memory System

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill description promotes memory, audit trails, and persistent storage but does not clearly warn users that their information may be retained across multiple stores and logged in git/audit artifacts. This undermines informed consent and increases privacy risk, particularly for sensitive personal preferences, identities, or interaction history.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill explicitly establishes broad natural-language retention and audit logging across several memory stores, creating durable records of user-supplied information. Even if intended for helpful continuity, the breadth of collection and persistence raises privacy and data-minimization concerns.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The examples and trigger rules instruct the agent to persist future user statements into semantic and core memory, normalizing storage of user attributes and preferences. Without stronger safeguards, this can accumulate sensitive behavioral profiles over time and persist them beyond user expectations.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The architecture prescribes append-only episode logs, a vault that never auto-decays, reflection archives, and audit files, all of which can preserve detailed interaction history indefinitely. Such durable multi-store retention increases exposure if the workspace is accessed by other agents, users, backups, or source-control systems.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The forget triggers include routine correction phrases such as 'never mind' and 'scratch that,' which may refer only to the current conversational turn rather than a request to alter stored memory. This ambiguity can lead to unintended deletion, archiving, or corruption of persistent records.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The guide introduces routine capture of 'significant conversations' into dialogue archives for later loading, which creates a direct data retention risk. Natural-language conversations often contain secrets, personal data, or contextual identifiers, and storing them in reusable memory files increases the chance of later disclosure, prompt contamination, or cross-session leakage.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · UPGRADE-1.0.7.md (reported line 64)May include surrounding context.

Manual Upgrade Steps

Step 1: Create New Directories

bash
cd ~/.openclaw/workspace  # or your workspace path

Static analysis

Detected: suspicious.prompt_injection_instructions

Prompt-injection style instruction pattern detected.

Warn
Code
suspicious.prompt_injection_instructions
Location
references/architecture.md:1009