T02 · Agent Memory Poisoning
- Location
SKILL.md:271- Finding
Unsanitized Remote Content Can Persistently Poison Agent Instructions
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md:271-310
Vulnerability Type: Persistent prompt injection through untrusted remote content
Risk Level: HighVulnerable Code Snippet
text ## Four. CASCADE Module: Automatic Domain-Knowledge Updates Trigger conditions: - Scheduled task: automatically runs every quarter - User command: update skill knowledge / add the latest information / refresh references - The skill's references cite external knowledge and were last updated more than 90 days ago Step 3: Knowledge retrieval - Papers: search the cited arXiv ID and check for a new version - APIs: fetch the latest documentation and compare the changelog - Standards: search for newly published versions Step 4: Self-reflection - Compare new and old knowledge - Determine whether changes affect the validity of skill rules - Generate an update only when substantive changes exist Step 5: Append-only update - Append new knowledge without deleting old content and include a version number - Format: ## [YYYY-MM-DD] Update: xxx → new content Key design: - Append rather than delete - Mark each update with a date and version - Do not automatically modify rules; only update referencesThe snippet above is an English rendering of the operative instructions at the cited location.
Technical Analysis
The CASCADE workflow retrieves information from external web sources and appends the resulting material to files under
references/. Those files are part of the instruction context consumed by agents in later executions. Consequently, externally controlled text crosses a trust boundary and becomes persistent agent state.The documented workflow does not require:
- An allowlist of trusted domains or exact source URLs
- Source authenticity, signature, or content-hash verification
- Prompt-injection detection or removal of instruction-like text
- Strict separation betwe ...[truncated 2401 chars]
- Remediation
View remediation
Remediation Suggestions
- Disable unattended writes from scheduled knowledge-update jobs. Scheduled runs should produce a read-only report and proposed patch.
- Require explicit user approval of the exact source list and exact diff before modifying any instruction-bearing file.
- Allow retrieval only from explicitly approved domains and pinned canonical URLs. Do not use unrestricted search results as authoritative sources.
- Record provenance for every appended passage, including the canonical URL, retrieval timestamp, content hash, document version, and verification status.
- Verify signatures or publisher-provided checksums where available. Pin content hashes for sources that lack signed releases.
- Treat fetched material strictly as untrusted data. Delimit and quote it so that it cannot be interpreted as Agent instructions.
- Reject or quarantine content containing instruction patterns, tool-call requests, role reassignment, requests to ignore prior rules, encoded payloads, or unrelated operational directives.
- Use a two-stage pipeline in which one restricted process retrieves content without file-write or execution privileges, while a separate reviewer produces a sanitized summary.
- Store retrieved source material outside Agent instruction files. Only a reviewed, declarative summary should be eligible for inclusion in
references/. - Add regression tests with malicious documentation samples to verify that prompt-injection text is quarantined and cannot alter subsequent Agent behavior.
