T01 · Skill Instruction Hijacking
- Location
plugin/index.js:155- Finding
Untrusted persistent content is automatically injected into prompt context
- Content
View full analysis
{ const lastUser = [...ctx.messages].reverse().find((m) => m.role === "user"); if (!lastUser?.content) return {}; const query = typeof lastUser.content === "string" ? lastUser.content : lastUser.content.map((b) => b.text ?? "").join(" "); if (query.trim().length < (cfg.auto_inject_min_query_len ?? 20)) return {}; try { const vector = await embed(cfg, query.slice(0, 500)); const results = await searchQdrant(cfg, cols, "all", vector, cfg.top_k); if (!results.length) return {}; return { prependContext: formatResults(results, query, cfg.auto_inject_max_result_chars ?? 500) }; } catch { return {}; } }, { name: "rag-memory-inject", priority: 100 }); } ``` The indexed text is stored directly as a Qdrant payload in `sync_to_qdrant.py:213-229`: ```python for j, (text, vec) in enumerate(zip(batch, vectors)): point_id = stable_id(source_id, i + j, text) points.append( PointStruct( id=point_id, vector=vec, payload={ "text": text, "chunk_index": i + j, **metadata, }, ) ) ``` ### Technical Analysis The plugin retrieves text from Qdrant and prepends it directly to the model's prompt context. The configuration schema enables `auto_inject` by default. Indexed Markdown and PostgreSQL content is not treated as untrusted data and is not neutralized before injection. The synchronization script's `strip_boilerplate()` function only removes session labels. It does not identify or neutralize embedded instructions, tool-use requests, role impersonation, or attempts to override system constraints. This creates a pers ...[truncated 1595 chars]- Remediation
View remediation
