T01 · Skill Instruction Hijacking
- Location
references/prompts.md:22- Finding
Untrusted paper content is inserted into generation prompts without instruction isolation
- Content
View full analysis
- Remediation
View remediation
{{cleaned_text}} ``` 3. Apply equivalent isolation to `summary_text`, `detailed_text`, and `contribution_text` during quality judgment because those values may already contain injected instructions. 4. Keep trusted instructions in a higher-priority message than document data where the runtime supports role-separated messages. 5. Add pre-generation detection for common injection patterns and surface a warning when suspicious directives are found. Detection should supplement, not replace, prompt isolation. 6. Restrict tool use during summarization so document text cannot induce network calls, file access, command execution, or unrelated actions. 7. Add adversarial tests using PDFs, DOCX files, and plain-text papers containing embedded role changes and “ignore previous instructions” payloads. ]]>
