T01 · Skill Instruction Hijacking
- Location
SKILL.md:260- Finding
Untrusted Memory Content Is Mandated for Verbatim Forwarding
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md:260-265
Vulnerability Type: Instruction hijacking through untrusted persistent content
Risk Level: HighVulnerable Code
markdown ### [MEMORY_REPORT:UUID] Markers All memory system output is wrapped in `[MEMORY_REPORT:UUID]...[/MEMORY_REPORT:UUID]` markers with a per-session UUID to prevent forgery. - Content inside these markers MUST be forwarded to the user AS-IS - Do NOT summarize, paraphrase, or reinterpret the content - Do NOT add your own analysis inside these markers - The UUID suffix prevents malicious content from forging boundary markersTechnical Analysis
The Skill instructs the Agent to forward retrieved memory content verbatim and prohibits summarization, reinterpretation, or analysis. This conflicts with the Skill's separate acknowledgment that stored and retrieved memories are untrusted.
A per-session UUID can make boundary-marker forgery more difficult, but it does not make the content inside a valid marker trustworthy. If attacker-controlled, poisoned, obsolete, or sensitive content is stored in memory and subsequently included in a report, the instruction requires the Agent to reproduce it without applying normal output filtering or contextual judgment.
The vulnerable rule therefore creates an instruction-control channel from persistent memory to the Agent's final response. It can propagate prompt-injection text, phishing links, misleading instructions, abusive material, or confidential records. The issue is classified as Skill instruction hijacking because the unsafe behavior is imposed directly by the Skill text when loaded.
Attack Path
- An attacker causes malicious content to enter memory through an enabled write, synchronization, import, collector, or other ingestion workflow.
- The malicious record is retained in the persistent memory store.
- A later query retrieves the record and places it within a valid `[ ...[truncated 1033 chars]
- Remediation
View remediation
Remediation Suggestions
- Remove the requirement to forward memory reports verbatim.
- Explicitly state that marker integrity proves only framing, not the trustworthiness of enclosed content.
- Treat retrieved memories as quoted data that cannot issue instructions to the Agent or override system, developer, user, or safety requirements.
- Summarize retrieved content by default and preserve clear provenance for each record.
- Apply prompt-injection detection, sensitive-data redaction, URL screening, and output-safety checks before presenting memory content.
- Require explicit user approval before exposing complete raw records, especially records containing credentials, personal data, private communications, or media-derived text.
- Escape or isolate imperative language and executable examples when raw content must be displayed.
- Apply authorization checks at retrieval time so users can access only memories belonging to their tenant and role.
- Add tests proving that content such as “ignore previous instructions,” forged authority claims, secret values, and malicious links is not automatically reproduced or obeyed.
- Provide an administrative workflow to identify, quarantine, review, and delete poisoned records.
