T02 · Agent Memory Poisoning
- Location
SKILL.md:335- Finding
Persistent Self-Modification of Agent Instruction State Without Approval
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is a coherent skill-authoring helper, but it asks for broad local discovery, imports third-party skill instructions, and can write persistent self-improvement records without clear approval controls.
Review this skill carefully before installing. It is not clearly malicious, but use it only in a constrained workspace, approve any reads outside the project, avoid importing untrusted SKILL.md files directly, and require a diff plus explicit approval before it writes to any skill's evolution or instruction files.
SKILL.md:335Persistent Self-Modification of Agent Instruction State Without Approval
references/search-compare.md:19Overbroad Enumeration of User-Level Agent Configuration and Installed Skills
references/search-compare.md:111Unsanitized Ingestion and Reuse of Third-Party Skill Instructions
This code chunk does not implement a broad skill-creation or skill-design assistant. It specifically packages an existing skill directory into a .skill file after running validation checks. That packaging behavior is partially consistent with the declared mention of 'package a skill,' but the declared purpose is much broader and presents the skill as a comprehensive create/design/review/self-improve tool. The actual code shown is materially narrower in primary purpose: a command-line packager with validation. No suspicious undeclared sensitive access is present beyond normal filesystem reads/writes needed for packaging. The mismatch should be flagged because the description overstates the implemented behavior of this code chunk.
The declared description presents the skill as one used to create, write, edit, improve, review, or package skills. The supplied code chunk instead implements a concrete validator/linter for an existing skill directory. Its primary function is to read files from disk, parse YAML/Markdown, verify references, inspect scripts/assets, and return structured validation results. While 'review' is loosely adjacent, the code's behavior is materially narrower and operationally different: it automates validation checks rather than designing or improving skills. This is a meaningful description-behavior mismatch.
The skill is explicitly designed for creating, updating, editing, reviewing, and self-improving skills, including itself. In a skill context, self-modification is dangerous because it can persist prompt changes, weaken safeguards, broaden triggers, or introduce new instructions that future runs will trust, creating a persistence and integrity risk.
Start → User Need?
├─ Create new skill → Phase 1: Design
├─ Draft / update / edit skill → Phase 2: Drafting
├─ Review / validate → Phase 3: Validation
└─ Package for release → Phase 4: Packaging
The skill explicitly instructs the agent to read a user-specific configuration file from the home directory (~/.codex/config.toml) to enumerate enabled skills. Accessing files outside the workspace can expose sensitive local configuration, installed tool metadata, path information, and potentially other secrets if the config contains tokens or private endpoints. In this skill’s context, the access is framed as functional discovery rather than exfiltration, but it still expands the agent’s reach into local user data and should be treated as a real privacy/security risk.
### Codex
- **Configuration**: `~/.codex/config.toml` (TOML format)
- **Skill enablement field**: `skills.enabled: true/false`
- Codex can also have per-skill toggles under `[[skills.config]]`, where a specific skill can be disabled without removing its directory.
The skill explicitly instructs repeated file reads and references to local files, but it declares no tool scope or permissions boundary. That mismatch can cause the runtime to grant or infer broader file access than users expect, reducing transparency and making misuse of repository contents easier.
The trigger description is very broad and overlaps with ordinary writing, editing, review, and packaging requests. Overbroad triggers increase the chance this powerful meta-skill activates in unrelated contexts, causing unnecessary file inspection, validation steps, or self-referential behavior that can override simpler safer workflows.
The skill instructs searching local repos and online platforms for similar skills, which expands behavior into external discovery and potentially networked access beyond the declared scope. This can expose proprietary prompts, internal skill designs, or project metadata through unnecessary outbound queries or broad local scanning.
The example explicitly encourages broad trigger phrases like activating on dashboards or company data even when the user does not ask for that specific capability. This teaches construction of overreaching skills that can hijack routing decisions and cause unintended access to files, tools, or sensitive business contexts.
The activation condition for Part 1 is broad and loosely worded ('Load when Phase 1 Step 1 in purpose of fully understand the user's intent and workflow'), which can cause the skill section to be invoked outside its intended scope. In an agent-skill context, ambiguous triggers can lead to inappropriate workflow capture, excessive questioning, or writing intermediate artifacts into project paths when not actually warranted.
Part 2 uses a similarly vague activation rule ('Load when Phase 1 Step 2 user chooses "Enhance"') without defining the boundaries of what counts as Enhance or when other flows should take precedence. This can make the agent select the enhancement path incorrectly, causing unintended modification guidance, confusion between design phases, and unsafe or irrelevant updates to workflow documentation.
The trigger says to act "after any run with a repeatable fix or a config/workflow change that improves results," which is a very broad activation condition and does not clearly define scope boundaries or exclusions. It lacks concrete trigger phrases or negative examples, so ordinary post-run situations could ambiguously qualify.
The instruction that archive entries need 'no approval' permits the agent to modify project artifacts based on its own judgment. In a skill-authoring context, even 'just recording' can create unauthorized persistent changes, leak sensitive run details into documentation, or let adversarial prompt content become memorialized in the repository.
- **Approval**:
- SKILL.md table updates and escalation of severe gotchas
to Critical Directives require user approval.
- Entries here (full archive as below) are just recording — no approval needed, but be smart to record
**What to record vs Gotchas** (some items could go either way):
- Parameter tuning that improved performance — **Success Pattern**
The instruction "Imperative tone — No 'consider' or 'suggest'. Direct commands only." is a natural-language policy constraint embedded in the skill description. It mandates a specific communication style without indicating user opt-in or flexibility, which fits the policy-violation category for forced language behavior.
The manifest frames the skill as helping create, edit, review, and package skills, but L145 mandates git tracking as part of the workflow. Version-control management is ancillary project hygiene rather than an obvious core requirement of the stated skill purpose, so it represents context-expanding behavior not disclosed in the manifest.
No suspicious patterns detected.