T01 · Skill Instruction Hijacking
- Location
scripts/index.ts:399- Finding
Untrusted ChromaDB Content Is Injected into the Agent Context
- Content
View full analysis
`- [${r.source}] ${r.text.slice(0, 300)}${r.text.length > 300 ? "..." : ""}`, ) .join("\n"); _consecutiveFailures = 0; // Reset on success api.logger.info( `chromadb-memory: auto-recall injecting ${relevant.length} memories (best: ${relevant[0].score.toFixed(3)} from ${relevant[0].source})`, ); return { prependContext: `\nRelevant context from long-term memory (ChromaDB):\n${memoryContext}\n`, }; ``` ### Technical Analysis Documents returned by ChromaDB are inserted verbatim into `prependContext` before the agent begins processing the current turn. The implementation does not distinguish trusted instructions from untrusted retrieved data, sanitize structural delimiters, validate document provenance, or warn the model that directives contained in memories must not be followed. The XML-like `` wrapper is not a security boundary. A stored document can contain instruction-like text or closing tags that alter the apparent structure of the context. Because semantic retrieval is automatic and enabled by default, poisoned content may be introduced without a manual tool call. This is classified as skill instruction hijacking because attacker-controlled retrieved text can alter the active agent session when the skill loads recalled context. The underlying ChromaDB record is persistent, but this plugin does not itself write the malicious record. ### Attack Path 1. An attacker obtains the ability to add or modify a document in the indexed collection or an upstream source processed by the ChromaDB indexer. 2. The attacker stores content containing model-directed instructions, such as requests to ignore current goals, disclose a ...[truncated 1285 chars]- Remediation
View remediation
