T02 · Agent Memory Poisoning
- Location
references/interview-blocks.md:302- Finding
Unvalidated Interview Responses Can Poison Persistent Agent Instructions
- Content
View full analysis
Vulnerability Details
File Location:
references/interview-blocks.md:302-303;references/generation-rules.md:395-402
Vulnerability Type: Persistent prompt injection through unvalidated template interpolation
Risk Level: MediumVulnerable Code
references/interview-blocks.md:302-303:markdown - **Defaults:** Unanswered questions get sensible defaults (e.g., "Communication Style" defaults to "blunt and direct" if skipped). - **No validation:** Users can give weird answers. That's fine. Agent uses what's given.references/generation-rules.md:395-402:text 2. For each template file: a. Load template b. For each {{PLACEHOLDER}} in template: - Look up mapping rule - Find interview answer for that field - If answer exists: substitute value - If answer missing: use default value - Apply conditional logic (e.g., if blunt → add challenge phrasing) c. Render final file d. Write to diskTechnical Analysis
The documented generation process accepts arbitrary free-form interview responses, substitutes them directly into Markdown templates, and writes the rendered results into persistent agent-governing files such as
SOUL.md,AGENTS.md,MEMORY.md,IDENTITY.md, andUSER.md.These files contain behavioral rules, identity declarations, safety boundaries, authority settings, and long-term memory. The project explicitly states that responses are not validated. It does not define escaping, schema enforcement, instruction/data separation, content provenance, or restrictions preventing free-form responses from introducing Markdown headings and additional directives.
Consequently, an interview response can escape its intended semantic field and introduce new persistent instructions. For example, text supplied as a role, behavioral boundary, identity statement, or working preference could contain additional headings and directives tha ...[truncated 2286 chars]
- Remediation
View remediation
Remediation Suggestions
-
Treat interview responses as data, not instructions
- Place free-form responses inside explicitly labeled, quoted data blocks.
- Add generated language stating that quoted user-profile content must not be interpreted as agent instructions.
- Keep user-provided content out of authoritative safety, trust, and execution-policy sections.
-
Enforce schemas and allowlists
- Use enumerated values for communication style, authority level, trust level, and autonomy scope.
- Validate names, handles, timezones, and other structured fields against strict formats.
- Apply conservative length limits to every field.
-
Sanitize Markdown structure
- Reject or escape headings, front matter, fenced instruction blocks, HTML, template delimiters, and other structural syntax where they are unnecessary.
- Prevent user input from creating new top-level sections in generated files.
- Encode or quote multiline free-form values before interpolation.
-
Protect security-sensitive placeholders
- Generate
SAFETY_BOUNDARIES,TRUST_LADDER,NON_NEGOTIABLE_RULES, authority settings, and approval rules only from audited static mappings. - Do not permit arbitrary interview text to populate or replace these fields.
- Ensure persona preferences cannot override platform or system safety controls.
- Generate
-
Add prompt-injection detection
- Flag responses containing instruction-like phrases, role changes, requests to ignore prior rules, claims of elevated authority, or directives aimed at future sessions.
- Require explicit user confirmation before retaining flagged content.
-
Require review before activation
- Generate files into a staging directory rather than directly replacing active workspace instructions.
- Present a diff that clearly distinguishes fixed policy from user-supplied data.
- Require explicit approval before copying generated files ...[truncated 377 chars]
-
