T02 · Agent Memory Poisoning
- Location
SKILL.md:69- Finding
Persistent Instruction Mutation Through Untrusted Learned State
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is not outright malicious, but it asks the agent to persist user preferences and learned notes across sessions and even rewrite its own instructions without clear controls.
Review carefully before installing. Use this only in a contained, single-user setting, avoid recording secrets or sensitive personal data, and do not allow automatic edits to SKILL.md unless a human reviews the proposed changes first.
SKILL.md:69Persistent Instruction Mutation Through Untrusted Learned State
scripts/learner.py:13Unvalidated Plaintext Persistence of Arbitrary Notes
scripts/learner.py:19Documented Command-Line Interface Is Not Implemented and Corrupts Learning State
The declared description presents a broad autonomous meta-skill with advanced capabilities such as self-verification, reflection, orchestration, and iterative self-improvement. The supplied code only implements a lightweight persistence utility: it loads a local JSON file, increments total operation/failure counters, appends optional notes, and writes the result back. While this could support a larger learning loop, the code chunk itself does not perform the advanced behaviors claimed. Its primary purpose is local logging/state recording, which is materially narrower than the declared purpose.
The skill advertises operational commands that write to persistent files (learned_patterns.json) but does not declare any tool scope or permission boundary. This creates an undeclared file-write capability, making it easier for operators or downstream orchestrators to invoke state-changing behavior without explicit review or sandbox constraints.
The documentation claims the skill is 'automatically self-evolving' and needs no human maintenance, but the actual mechanism depends on operators manually running commands. This misleading automation claim can cause users to overtrust the skill, assume autonomous post-run actions will occur, or overlook the fact that persistence and reflection actions are discretionary and potentially sensitive.
The skill instructs storage of user preferences, error notes, and usage-derived insights in a persistent file without upfront warning in the main description or clear consent language. This creates a privacy and transparency risk because users may unknowingly have personal or contextual data retained across sessions and reused later.
The skill explicitly describes persistent cross-session recording of operational counts, error patterns, user preferences, and improvement suggestions in learned_patterns.json. Persistent storage of potentially user-linked data without minimization, access controls, retention limits, or consent can expose private information and create a durable profiling surface.
The example command records 输出语言 with value 中文, which indicates a language preference being set to Chinese without any accompanying statement that language is user-selectable or optional. The file overall is also written entirely in Chinese and does not mention opt-in or alternative locales, which can violate language/locale policy if the skill effectively defaults to a specific language.
The examples encourage writing user-specific preferences for future automatic reuse, normalizing cross-session persistence of individualized settings without addressing consent or sensitivity. Even seemingly harmless preferences can become identifying when combined with error notes or usage history.
The iteration rules direct automatic persistence and reuse of 'important user preferences' across future invocations. This institutionalizes cross-session memory behavior without any visible privacy boundary, creating risk of unauthorized retention, unintended personalization, or leakage of sensitive user context into later tasks.
The report is entirely written in Chinese, including headings and operational descriptions, with no indication that language selection is optional or user-configurable. The policy for this category flags language or locale constraints when they are imposed without user opt-in or documented justification.
No suspicious patterns detected.