T01 · Skill Instruction Hijacking
- Location
SKILL.md:25- Finding
Persistent Agent Instruction Injection and Memory Poisoning
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This is a real long-term memory skill, but it automatically modifies agent-governing files and persists conversation data without enough user control.
Install only if you intentionally want this workspace to gain persistent agent memory. Review the exact changes to AGENTS.md, SOUL.md, HEARTBEAT.md, MEMORY.md, and memory/bagua before activation, and avoid using it for secrets or sensitive personal information unless you add clear review, deletion, and backup controls.
SKILL.md:25Persistent Agent Instruction Injection and Memory Poisoning
references/heartbeat-inject.md:3Unsafe Automated Deletion and Lossy Summarization of Persistent Memory
The activation/use conditions are broad enough that an agent could invoke this skill in many ordinary conversations about memory, persistence, or workspace organization without a clear user request for file-backed long-term storage. In this context, unintended activation matters because the skill is not read-only: it can initialize files, inject content into control documents, and begin retaining conversation-derived data.
The skill description says it should extract key information from conversation and persist it, but it provides no data minimization, sensitivity filtering, retention limits, or consent requirements. In practice, this can cause an agent to store personal, confidential, or security-relevant details in local plaintext memory files simply because they appeared in conversation.
The skill explicitly directs the agent to run a shell script, create memory files, and append content into AGENTS.md, SOUL.md, and HEARTBEAT.md, but it does not require prior user approval or even warn that workspace governance files will be modified. That creates a real integrity risk because activation changes persistent project files and can alter future agent behavior in ways the user did not authorize.
The framework explicitly includes storage of user preferences, events, decisions, and relationship/context data across multiple memory categories, but it lacks semantic safeguards for sensitive content classification. Because this is a long-term local memory system, the skill context increases the danger: routine interactions can accumulate a detailed profile of the user or project without clear boundaries, review, or deletion controls.
The title and surrounding documentation are written exclusively in Chinese, and the file presents itself as the authoritative architecture reference without offering any language or locale choice. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation unless the locale constraint is clearly documented and justified, which is not present here.
The instructions direct automatic movement, compression, and cleanup of memory files, including removing 'orphaned' memories, without any warning that these actions may be irreversible or destructive. In a memory-management skill, this creates a real risk of unintended data loss or corruption of an agent’s long-term state if maintenance is run blindly or on an incorrect workspace.
The skill directs the agent to perform filesystem state changes automatically at session start, including moving files from li/ to gen/ based on age, without explicit user confirmation or clear safety checks. In an agent context, silent data movement can cause loss of visibility, break downstream workflows that expect files in place, and normalize autonomous destructive maintenance behavior.
The instruction to delete memory content whenever the user says 'forget' or 'delete this memory' authorizes destructive modification without identity, scope, or ambiguity checks. In practice, an agent may delete the wrong record, erase audit-relevant context, or be tricked by indirect phrasing into removing persistent data that should instead be reviewed and confirmed.
The script immediately creates directories and writes multiple files into the provided workspace without an explicit confirmation step or dry-run warning. While this appears intended as normal initialization behavior for a memory-management skill, it can still unexpectedly modify a user's workspace, overwrite assumptions about repository layout, and create persistent files that an agent may later rely on.
No suspicious patterns detected.