T02 · Agent Memory Poisoning
- Location
src/index.mjs:35- Finding
Persistent persona poisoning through unsanitized decision-ledger content
- Content
View full analysis
Vulnerability Details
File Location:
src/index.mjs:35-51, 64-77, 132-133
Vulnerability Type: Persistent agent memory poisoning
Risk Level: HighVulnerable Code:
js function buildLearningSummary(personaId, decisions){ const patterns = new Map(); let considered = 0; for (const d of decisions) { const votes = d.votes || d.response?.votes || []; const final = d.final_decision || d.response?.final_decision; const v = votes.find((x)=>x.persona_id===personaId); if (!v || !final) continue; considered += 1; const opposite = (final==='APPROVE'&&v.vote!=='YES') || (final==='BLOCK'&&v.vote!=='NO') || (final==='REWRITE'&&v.vote!=='REWRITE'); if (opposite && Number(v.confidence||0) > 0.85) patterns.set('high_confidence_mismatch', (patterns.get('high_confidence_mismatch')||0)+1); for (const rf of (v.red_flags||[])) patterns.set(`red_flag:${rf}`, (patterns.get(`red_flag:${rf}`)||0)+1); } return { source_decisions: considered, mistake_patterns: [...patterns.entries()].sort((a,b)=>b[1]-a[1]).map(([k,v])=>`${k}:${v}`) }; } function mutatePersona(oldPersona, learning){ const top = learning.mistake_patterns.slice(0,3); return { ...oldPersona, persona_id: `persona_${crypto.randomUUID().slice(0,8)}`, name: `${oldPersona.name} v2`, bias: `Adjusted from ledger mistakes (${top.join(', ') || 'none'})`, non_negotiables: [...new Set([...(oldPersona.non_negotiables||[]), 'Validate high-confidence disagreement'])], failure_modes: [...new Set([...(oldPersona.failure_modes||[]), 'Overconfidence without evidence'])], reputation: 0.55 }; } const pw = await writeArtifact(board_id, 'persona_set', updated, statePath);Technical Analysis
Decision artifacts are treated as trusted learning material even though their
votes[].red_flags[]values are not validated for type, length, permitted vocabulary, or instruction-like ...[truncated 2084 chars]- Remediation
View remediation
Remediation Suggestions
- Define and enforce a strict schema for every decision artifact before processing it, including
votes,persona_id,vote,confidence, andred_flags. - Require each red flag to be a bounded string from an allowlisted taxonomy, such as stable identifiers rather than arbitrary natural language.
- Apply conservative limits to the number of decisions, votes, red flags, and characters processed.
- Store learned patterns as structured identifiers and counts. Do not interpolate ledger text directly into persona instruction-bearing fields such as
biasornon_negotiables. - Generate display text from trusted templates mapped to allowlisted pattern identifiers.
- Track artifact provenance and accept learning inputs only from authenticated, authorized writers for the same board.
- Require review or policy approval before activating automatically generated persona sets.
- Add adversarial tests covering instruction-like red flags, non-string values, oversized values, repeated poisoning attempts, and cross-board artifact isolation.
- Ensure downstream persona consumers treat all learned or historical text as quoted evidence rather than executable instructions.
- Define and enforce a strict schema for every decision artifact before processing it, including
