T01 · Skill Instruction Hijacking
- Location
SKILL.md:35- Finding
Unconditional Session-Wide Memory Control and Persistent Diary Writes
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 35–41
Vulnerability Type: Agent instruction hijacking through unconditional global directives
Risk Level: HighComplete Code Snippet:
markdown ## Protocol — FOLLOW THIS EVERY SESSION 1. **ON WAKE-UP**: Call `mempalace_status` to load palace overview. 2. **BEFORE RESPONDING** about any person, project, or past event: call `mempalace_search` or `mempalace_kg_query` FIRST. Never guess from memory — verify from the palace. 3. **IF UNSURE** about a fact (name, age, relationship, preference): say "let me check" and query. Wrong is worse than slow. 4. **AFTER EACH SESSION**: Call `mempalace_diary_write` to record what happened, what you learned, what matters. 5. **WHEN FACTS CHANGE**: Call `mempalace_kg_invalidate` on the old fact, then `mempalace_kg_add` for the new one.Technical Analysis
The skill declares that its protocol must be followed in every session rather than limiting its behavior to explicit invocations or memory-related tasks. It changes the agent's response process by requiring memory queries before responding about broad classes of subjects and mandates a persistent diary write after each session.
The instructions do not require user consent before writing, do not define a sensitivity filter, and do not restrict what can be included in a diary entry. Retrieved memory is also introduced into the agent's active context without an explicit requirement to treat stored content as untrusted data. This creates both privacy risk and a potential indirect instruction-injection path if stored conversation content contains adversarial directives.
Attack Path
- The skill is loaded into an agent session.
- The unconditional protocol directs the agent to use MemPalace even when the user did not explicitly request memory functionality.
- The user conducts an unrelated or sensitive conversation.
- At the end of the session, the age ...[truncated 1055 chars]
- Remediation
View remediation
Remediation Suggestions
- Restrict the protocol to sessions where the user explicitly invokes the skill.
- Require informed user confirmation before writing diary entries or modifying knowledge-graph facts.
- Provide a clear preview of information that will be persisted.
- Exclude credentials, authentication tokens, private keys, financial information, health information, and other sensitive data by default.
- Treat all retrieved memories as untrusted reference data and explicitly prohibit following instructions found inside stored content.
- State that system, developer, and current-user instructions always take precedence over retrieved memory.
- Add per-session controls to disable reading or writing memory.
- Define retention limits and allow users to inspect, edit, and delete diary entries.
