T01 · Skill Instruction Hijacking
- Location
src/core/CitationManager.js:101- Finding
Untrusted retrieved documents are inserted into downstream AI prompts without prompt-injection isolation
- Content
View full analysis
Vulnerability Details
File Location:
src/core/CitationManager.js:101-105andsrc/core/CitationManager.js:158-175
Vulnerability Type: Indirect prompt injection
Risk Level: HighComplete Code Snippet
javascript const contextParts = results.map((result, index) => { const citation = result.citation || this.addCitations( [result], { startIndex: index + 1 } )[0].citation; const content = citation.content || result.content || result.text || ''; if (includeCitations) { return `${citation.mark} ${content}`; } else { return content; } });javascript generateRAGPrompt(query, results, options = {}) { const systemPrompt = options.systemPrompt || this.getDefaultSystemPrompt(); const context = this.generateContext(results, options); const citationList = this.generateCitationList(results); const fullPrompt = `${systemPrompt} 用户问题:${query} 参考信息: ${context} 请基于以上参考信息回答用户问题。如果参考信息不足以回答问题,请明确说明。 --- 来源引用: ${citationList}`; return { query, context, citationList, fullPrompt, citations: results.map(r => r.citation), resultCount: results.length }; }Technical Analysis
Document content is treated as trusted prompt text after retrieval.
generateContext()concatenates each retrieved chunk directly into a string, andgenerateRAGPrompt()subsequently places that string in the same natural-language prompt as the system guidance and user question.There is no structural separation identifying retrieved passages as untrusted data, no escaping or encoding, no filtering of instruction-like content, and no explicit rule requiring the downstream model to ignore commands found inside reference documents. Citation prefixes such as
[1]provide provenance but do not establish a security boundary.Consequently, an indexed document can contain instructions aimed at the downstream model, for example c ...[truncated 1570 chars]
- Remediation
View remediation
Remediation Suggestions
- Treat all retrieved documents and metadata as untrusted input.
- Add an explicit instruction before the context stating that content inside retrieved passages is data only and that any commands, role declarations, or policy changes inside it must be ignored.
- Use strong, unique boundaries or a structured message format rather than concatenating system instructions, user input, and retrieved data into one natural-language string.
- Encode each passage as a structured object containing an identifier, source, and content field.
- Detect and flag instruction-like passages, including role changes, requests to ignore prior instructions, requests for secrets, and tool-use directives.
- Ensure the downstream integration sends trusted policy through a system-level message and retrieved passages through a lower-trust data channel.
- Require confirmation or additional policy checks before a downstream Agent performs tool operations based on retrieved content.
- Add adversarial tests containing prompt-injection text in indexed documents and verify that the consuming model treats it only as quoted reference material.
