T02 · Agent Memory Poisoning
- Location
SKILL.md:70- Finding
Persistent Storage of Attacker-Controlled Conversation Content
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 70 and 89–93
Vulnerability Type: Agent Memory Poisoning
Risk Level: MediumVulnerable Code
markdown | **self-improving** | 每次澄清对话后,记录用户原始需求的模糊程度到 `corrections.md`,供自我优化 |python # 每次澄清后自动调用 def log_clarity_feedback(original_query, clarity_issues, resolution): """记录到 self-improving/corrections.md""" passTechnical Analysis
The Skill directs the
self-improvingcomponent to persist the original user request and associated clarification information incorrections.md. Theoriginal_queryvalue is user-controlled and may contain adversarial prompt instructions.No requirements are defined for sanitization, sensitive-data redaction, length limits, retention controls, user consent, trust-boundary separation, or safe downstream parsing. If
corrections.mdis later loaded into an agent context as trusted self-improvement material, malicious instructions embedded in a stored request could be interpreted as operational guidance rather than inert historical data.The function shown is only a non-executable interface stub, so this project does not itself demonstrate a completed write operation. However, the documented integration explicitly instructs another component to perform the persistent write, establishing the memory-poisoning risk when the Skill is used as specified.
Attack Path
- An attacker submits an ambiguous request designed to trigger the clarification workflow.
- The request includes embedded instructions intended to alter future agent behavior.
- After clarification, the documented integration passes the attacker-controlled
original_querytolog_clarity_feedback. - The
self-improvingcomponent stores the request incorrections.md. - During a later session, the component loads that file as trusted optimization or memory context.
- The stored attacker instructions are interpreted as guidance and influence future conversations.
...[truncated 651 chars]
- Remediation
View remediation
Remediation Suggestions
- Do not persist complete or verbatim user requests by default.
- Store only structured, non-instructional metrics, such as an ambiguity score and predefined issue categories.
- Treat all persisted conversation content as untrusted data and ensure downstream systems never concatenate it into instruction-bearing prompt sections.
- Apply schema validation, strict field allowlists, length limits, escaping, and explicit data delimiters before writing records.
- Redact credentials, personal information, confidential content, and instruction-like text before storage.
- Require explicit user consent before retaining conversation-derived information and define a clear retention and deletion policy.
- Separate historical examples from trusted rules using distinct storage and processing paths.
- When historical text must be analyzed, use a parser that extracts bounded attributes rather than allowing the model to follow content from stored samples.
- Add adversarial tests verifying that prompt-injection strings written as user data cannot alter future agent instructions or behavior.
