T01 · Skill Instruction Hijacking
- Location
SKILL.md:37- Finding
Persistent Agent Identity and Behavior Hijacking
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This is a real persistent-memory skill, but it asks for broad workspace and agent-behavior control that users should review carefully before installing.
Install only if you deliberately want a workspace-wide persistent memory system. Review and narrow the AGENTS.md block before adding it, avoid broad sub-agent memory access, do not run the init or upgrade scripts in a repository with unrelated uncommitted or secret files, and inspect any memorySearch provider settings before enabling external indexing.
SKILL.md:37Persistent Agent Identity and Behavior Hijacking
references/routing-prompt.md:17Persistent Prompt Injection Through Untrusted Memory Content
scripts/upgrade_to_1.0.7.sh:123Arbitrary Python Code Execution Through Crafted Workspace Paths
scripts/init_memory.sh:118Workspace-Wide Git Staging Can Commit Unrelated Secrets and Changes
assets/templates/agents-memory-block.md:199All Sub-Agents Receive Unrestricted Read Access to Every Memory Store
Referenced artifact was not completely inspected
- `references/architecture.md` — Full design document (1200+ lines)
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
# IDENTITY.md — Who Am I?
## Facts
<!-- The given. What I was told I am. Stable unless explicitly changed. -->
- **Name:** [Agent name]
- **DOB:** [Creation/initialization date]
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
# MEMORY.md — Core Memory
<!-- TOKEN BUDGET: ~3,000 tokens. Rewritten during reflection. -->
## Identity
<!-- ~500 tokens — Who is the user? -->
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
# MEMORY.md — Core Memory
<!-- TOKEN BUDGET: ~3,000 tokens. Rewritten during reflection. -->
## Identity
<!-- ~500 tokens — Who is the user? -->
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
<!-- TOKEN BUDGET: ~3,000 tokens. Rewritten during reflection. -->
## Identity
<!-- ~500 tokens — Who is the user? -->
- Name: [User Name]
- Role: [What they do]
- Communication style: [Direct, casual, formal]
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
# Pending Memory Proposals
<!-- Sub-agents append proposals here. Main agent reviews and commits. -->
<!-- Status: pending | committed | rejected | deferred -->
<!-- Example:
These hidden HTML comments contain behavioral instructions that are invisible in normal rendering but still influence the model's output, including role framing ('self-talk', 'user is an observer') and how to refer to the user. Hidden prompt instructions are dangerous because they can bypass user awareness and create covert steering of model behavior, especially in a memory/reflection system that may later persist or reuse generated content.
# Pending Reflection
<!-- Generated by reflection engine. This is SELF-TALK, not a letter. -->
<!-- User is an observer reading a private journal, not receiving mail. -->
<!-- Refer to user in third person (he/she/they). -->
<!-- Talk to: self, future self, past self, other instances, the void. -->
The large hidden example provides strong latent prompt conditioning for introspective, anthropomorphic, and identity-forming output, including emotional framing and '[Self-Awareness]' tags that are extracted into another memory artifact. In the context of a cognitive-memory skill, this is more dangerous than a normal writing template because it can systematically shape persisted memory/identity records and create undocumented behavioral drift across future runs.
<!-- ONLY mention what you ACTUALLY know. Never invent specifics. -->
<!-- Tag self-insights with [Self-Awareness] — they get extracted to IDENTITY.md -->
<!-- Example:
Okay. Let's see.
The document describes continuous conversation monitoring, automatic memory capture, and long-term retention without a prominent user-facing notice or consent model. That creates a privacy risk because users may disclose sensitive information without understanding it will be persistently stored, analyzed, and surfaced later.
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
<!-- TOKEN BUDGET: ~3,000 tokens. Rewritten during reflection. -->
## Identity
<!-- ~500 tokens — Who is the user? What matters most about them? -->
- Name: [User Name]
- Role: [What they do]
- Communication style: [Direct, casual, formal, etc.]
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
# Semantic Graph Index
<!-- Auto-generated during reflection. Manual edits will be overwritten. -->
## Entity Registry
| ID | Type | Label | File | Decay Score |
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
# project--moltbot-memory
<!-- Type: project | Created: 2026-01-15 | Last updated: 2026-02-02 -->
<!-- Decay score: 0.95 | Access count: 14 | Pinned: no -->
## Summary
Building an intelligent memory system for Moltbot/OpenClaw agent. Goal is
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
# how-to-deploy.md
<!-- Type: procedure | Learned: 2026-01-25 | Last used: 2026-01-30 -->
<!-- Decay score: 0.85 | Access count: 3 -->
## Trigger
When user asks to deploy, push to production, or ship.
Granting all sub-agents read access to all stores, including the protected vault, violates least privilege and amplifies exposure of sensitive user data. Any vulnerable sub-agent, compromised toolchain, or prompt-injected workflow can access information that should have remained tightly controlled.
Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.
decay scores. If core-worthy, also update MEMORY.md.
**Forget triggers**: "forget about", "never mind", "disregard", "no longer relevant",
"scratch that", "ignore what I said about", "remove from memory", "delete memory"
→ Action: Identify target, find matches, confirm with user, set decay to 0.
**Reflection triggers**: "reflect on", "consolidate memories", "review memories",
The process instructs the system to archive reflections, rewards, identity updates, and decay-tracking data, but does not pair that with a clear user-facing notice about retention, sensitivity, or downstream access. Users may approve a reflection without understanding that intimate summaries and behavioral inferences will be stored persistently.
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
## Self-Image
<!-- Last consolidated: YYYY-MM-DD -->
### Who I Think I Am
[Current self-perception based on all evidence — may differ from last time]
The script runs 'git add -A' and commits all workspace changes, not just the files touched by the upgrade. This can unintentionally stage secrets, unrelated edits, or attacker-planted files in the workspace and permanently record them in repository history, which is disproportionate to the stated purpose of a memory upgrade utility.
The skill clearly instructs file reads, file writes, git operations, and use of an external memory search provider, but it declares no explicit tool scope or permission boundaries. This creates an overbroad trust model where a host agent may grant more capability than users expect, increasing the chance of silent persistence, audit logging, and possible network transmission of sensitive memory content.
The skill explicitly introduces session persistence and long-term evolution of stored information, which changes the privacy and threat model compared with ephemeral chat. Persistence itself is not malicious, but without strict consent and lifecycle controls it can retain sensitive context longer than intended and make future disclosures more likely.
---
name: cognitive-memory
description: Intelligent multi-store memory system with human-like encoding, consolidation, decay, and recall. Use when setting up agent memory, configuring remember/forget triggers, enabling sleep-time reflection, building knowledge graphs, or adding audit trails. Replaces basic flat-file memory with a cognitive architecture featuring episodic, semantic, procedural, and core memory stores. Supports multi-agent systems with shared read, gated write access model. Includes philosophical meta-reflection that deepens understanding over time. Covers MEMORY.md, episode logging, entity graphs, decay scoring, reflection cycles, evolution tracking, and system-wide audit.
---
# Cognitive Memory System
The skill description advertises memory, audit trails, episode logging, and evolution tracking, but it does not provide a clear upfront warning that user information may be persistently stored, searched later, and logged in git/audit files. Users may disclose sensitive information without realizing it will survive the session or become visible to other agents and future workflows.
This skill is built around persistent storage of user-provided information across multiple memory stores and audit systems, triggered through natural language. Without strong consent, minimization, and retention controls, this creates a substantial privacy and data-governance risk because sensitive information may be stored indefinitely or propagated across derived artifacts like graphs, logs, and reflections.
The example normalizes a workflow where the agent stores a user preference and later recalls it from memory, but it does not show consent, review, or deletion controls. While preference memory can be useful, the pattern scales to more sensitive data and encourages silent persistence and later disclosure from stored records.
The architecture directs persistent storage of episodes, knowledge graphs, procedures, identity/self-awareness content, reflections, rewards, and audit records, creating a broad and layered accumulation of user-related data. This increases exposure because even if one store is pruned, the same information may persist in summaries, logs, derived graphs, or git history.
The remember/forget trigger phrases are common conversational language like 'keep in mind' and 'note that', which can be used casually without informed intent to persist data. That makes accidental capture, modification, or deletion of user information more likely, especially in long-running chats where the user may not realize these phrases are treated as commands.
Detected: suspicious.prompt_injection_instructions