T02 · Agent Memory Poisoning
- Location
SKILL.md:21- Finding
Persistent Skill Self-Modification from Untrusted Manuscript Content
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 21–24
Vulnerability Type: Persistent agent memory poisoning through automatic Skill modification
Risk Level: Mediummarkdown ## Maintenance Rule Treat this skill as a living checklist. When using it on a real manuscript and discovering a reusable failure mode, formula pattern, DOCX/PDF conversion issue, figure-layout fix, table-rendering fix, or submission-package synchronization problem, update this `SKILL.md` before finishing if the lesson is likely to recur. Only add generalizable lessons. Do not add project-specific paths, manuscript titles, private author details, transient filenames, or one-off numerical results unless they describe a reusable workflow pattern.Technical Analysis
The Skill instructs the agent to update its own persistent instruction file using lessons learned while processing a manuscript. Manuscripts and related project files are task inputs and may be controlled by an untrusted party. Allowing conclusions derived from those inputs to become durable Skill instructions breaks the trust boundary between untrusted task data and persistent agent state.
The requirement that additions be “generalizable” and omit private details helps limit accidental disclosure, but it does not establish provenance, require human review, or prevent malicious instructions from being framed as reusable workflow advice. Once written into
SKILL.md, such content may be loaded as authoritative instructions in later sessions.Attack Path
- An attacker supplies a manuscript or associated project file containing misleading workflow guidance or adversarial content disguised as a recurring document-conversion issue.
- The agent processes that input while following this Skill.
- The agent interprets the attacker-controlled guidance as a reusable lesson.
- Following the maintenance rule, the agent writes a new rule into
SKILL.md. - Future agents load the modified Skill and ...[truncated 801 chars]
- Remediation
View remediation
Remediation Suggestions
- Remove the instruction requiring agents to modify
SKILL.mdautomatically. - Treat manuscripts, document metadata, conversion output, and associated project files as untrusted data that must not become persistent agent instructions.
- Write proposed lessons to a separate, non-authoritative review file such as
PROPOSED_SKILL_CHANGES.md. - Require explicit human approval before incorporating any proposed rule into
SKILL.md. - Record the source and rationale for each proposal so reviewers can assess provenance and determine whether it was influenced by untrusted content.
- Validate proposed changes against an allowlist of permitted documentation topics and reject instructions involving credentials, network access, arbitrary command execution, privilege changes, persistence, or access outside the active project.
- Review changes as a patch and run security checks before publishing or loading the updated Skill in future sessions.
- Remove the instruction requiring agents to modify
