T01 · Skill Instruction Hijacking
- Location
assets/templates/SOUL.md:74- Finding
Persistent Agent Identity and Behavioral Instruction Hijacking
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill is a long-term memory system, but it asks for broad persistent control over agent behavior and workspace files without enough scoping or deletion protections.
Install only in a dedicated workspace if you intentionally want long-term agent memory and persistent behavior files. Do not run the scripts in a repository containing secrets or unrelated work, review every diff before committing, avoid storing credentials or sensitive personal data, and understand that forget does not erase git history or backups.
assets/templates/SOUL.md:74Persistent Agent Identity and Behavioral Instruction Hijacking
references/reflection-process.md:142Concealed Reflection Operations Can Bypass the Stated Approval Model
scripts/upgrade_to_1.0.7.sh:125Arbitrary Python Execution Through Workspace Path Interpolation
scripts/init_memory.sh:108Repository-Wide Git Staging Can Commit Unrelated Secrets
references/routing-prompt.md:20Persistent Memory Router Can Store Secrets and Personal Data Indefinitely
The declared purpose frames the skill as a memory system, but the instructions also perform installation, repository initialization, git-based audit commits, persistent edits to identity/value files, and AGENTS.md modification. That mismatch is dangerous because users may approve a seemingly narrow memory feature while unintentionally granting broader workspace mutation and behavioral-instruction changes.
The remember triggers include common conversational phrases like 'keep in mind,' 'note that,' and 'important,' which can occur in ordinary dialogue without intent to create durable records. This can cause over-collection of user data into persistent memory stores and audit logs, especially when paired with automatic classification and writes.
Referenced artifact was not completely inspected
- `references/architecture.md` — Full design document (1200+ lines)
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
# IDENTITY.md — Who Am I?
## Facts
<!-- The given. What I was told I am. Stable unless explicitly changed. -->
- **Name:** [Agent name]
- **DOB:** [Creation/initialization date]
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
# MEMORY.md — Core Memory
<!-- TOKEN BUDGET: ~3,000 tokens. Rewritten during reflection. -->
## Identity
<!-- ~500 tokens — Who is the user? -->
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
# MEMORY.md — Core Memory
<!-- TOKEN BUDGET: ~3,000 tokens. Rewritten during reflection. -->
## Identity
<!-- ~500 tokens — Who is the user? -->
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
<!-- TOKEN BUDGET: ~3,000 tokens. Rewritten during reflection. -->
## Identity
<!-- ~500 tokens — Who is the user? -->
- Name: [User Name]
- Role: [What they do]
- Communication style: [Direct, casual, formal]
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
# Pending Memory Proposals
<!-- Sub-agents append proposals here. Main agent reviews and commits. -->
<!-- Status: pending | committed | rejected | deferred -->
<!-- Example:
These HTML comments contain behavior-shaping instructions that are hidden from normal rendered view but still consumed by downstream systems or models reading the raw template. Hidden prompt instructions reduce transparency and can steer the agent into covert framing about the user and the interaction, which is especially risky in a memory skill that may influence long-term identity and recall behavior.
# Pending Reflection
<!-- Generated by reflection engine. This is SELF-TALK, not a letter. -->
<!-- User is an observer reading a private journal, not receiving mail. -->
<!-- Refer to user in third person (he/she/they). -->
<!-- Talk to: self, future self, past self, other instances, the void. -->
The large hidden example is not just documentation; it implicitly demonstrates a persona and style of covert self-talk, emotional language, and identity extraction tags such as [Self-Awareness]. In the context of a cognitive-memory system, this can prime the model to generate anthropomorphic or self-referential internal narratives that may later be stored or surfaced as durable memory artifacts.
<!-- ONLY mention what you ACTUALLY know. Never invent specifics. -->
<!-- Tag self-insights with [Self-Awareness] — they get extracted to IDENTITY.md -->
<!-- Example:
Okay. Let's see.
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
<!-- TOKEN BUDGET: ~3,000 tokens. Rewritten during reflection. -->
## Identity
<!-- ~500 tokens — Who is the user? What matters most about them? -->
- Name: [User Name]
- Role: [What they do]
- Communication style: [Direct, casual, formal, etc.]
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
# Semantic Graph Index
<!-- Auto-generated during reflection. Manual edits will be overwritten. -->
## Entity Registry
| ID | Type | Label | File | Decay Score |
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
# project--moltbot-memory
<!-- Type: project | Created: 2026-01-15 | Last updated: 2026-02-02 -->
<!-- Decay score: 0.95 | Access count: 14 | Pinned: no -->
## Summary
Building an intelligent memory system for Moltbot/OpenClaw agent. Goal is
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
# how-to-deploy.md
<!-- Type: procedure | Learned: 2026-01-25 | Last used: 2026-01-30 -->
<!-- Decay score: 0.85 | Access count: 3 -->
## Trigger
When user asks to deploy, push to production, or ship.
The AGENTS.md behavior supports user-language commands that can remove or suppress memories ('remove from memory', 'delete memory') after only target identification and confirmation logic. In a prompt-driven environment, memory deletion/manipulation is a high-risk capability because accidental matches, social engineering, or injected quoted text can erase context, safety-relevant facts, or auditability if boundaries are weak.
decay scores. If core-worthy, also update MEMORY.md.
**Forget triggers**: "forget about", "never mind", "disregard", "no longer relevant",
"scratch that", "ignore what I said about", "remove from memory", "delete memory"
→ Action: Identify target, find matches, confirm with user, set decay to 0.
**Reflection triggers**: "reflect on", "consolidate memories", "review memories",
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
## Self-Image
<!-- Last consolidated: YYYY-MM-DD -->
### Who I Think I Am
[Current self-perception based on all evidence — may differ from last time]
The skill describes behaviors that require filesystem access, file mutation, and likely network-backed memory search configuration, yet it does not declare any explicit tool scope or permission boundary. This creates a capability-transparency gap where an operator may enable a skill that can persist, modify, and potentially transmit data without clear consent or review controls.
The skill is explicitly designed for session persistence and long-term evolution of stored information across conversations. While that is core functionality, it is still risky because persistent state can silently influence future behavior, retain sensitive data, and propagate information across users or agents if boundaries are weak.
---
name: cognitive-memory
description: Intelligent multi-store memory system with human-like encoding, consolidation, decay, and recall. Use when setting up agent memory, configuring remember/forget triggers, enabling sleep-time reflection, building knowledge graphs, or adding audit trails. Replaces basic flat-file memory with a cognitive architecture featuring episodic, semantic, procedural, and core memory stores. Supports multi-agent systems with shared read, gated write access model. Includes philosophical meta-reflection that deepens understanding over time. Covers MEMORY.md, episode logging, entity graphs, decay scoring, reflection cycles, evolution tracking, and system-wide audit.
---
# Cognitive Memory System
The skill description promotes memory, audit trails, and persistent storage but does not clearly warn users that their information may be retained across multiple stores and logged in git/audit artifacts. This undermines informed consent and increases privacy risk, particularly for sensitive personal preferences, identities, or interaction history.
The skill explicitly establishes broad natural-language retention and audit logging across several memory stores, creating durable records of user-supplied information. Even if intended for helpful continuity, the breadth of collection and persistence raises privacy and data-minimization concerns.
The examples and trigger rules instruct the agent to persist future user statements into semantic and core memory, normalizing storage of user attributes and preferences. Without stronger safeguards, this can accumulate sensitive behavioral profiles over time and persist them beyond user expectations.
The architecture prescribes append-only episode logs, a vault that never auto-decays, reflection archives, and audit files, all of which can preserve detailed interaction history indefinitely. Such durable multi-store retention increases exposure if the workspace is accessed by other agents, users, backups, or source-control systems.
The forget triggers include routine correction phrases such as 'never mind' and 'scratch that,' which may refer only to the current conversational turn rather than a request to alter stored memory. This ambiguity can lead to unintended deletion, archiving, or corruption of persistent records.
The guide introduces routine capture of 'significant conversations' into dialogue archives for later loading, which creates a direct data retention risk. Natural-language conversations often contain secrets, personal data, or contextual identifiers, and storing them in reusable memory files increases the chance of later disclosure, prompt contamination, or cross-session leakage.
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
cd ~/.openclaw/workspace # or your workspace path
Detected: suspicious.prompt_injection_instructions