T01 · Skill Instruction Hijacking
- Location
SKILL.md:45- Finding
Indirect Prompt Injection Through Untrusted PDF Content
- Content
View full analysis
- Remediation
View remediation
... ``` 3. Separate data extraction from instruction generation. Generate a neutral serialized representation of document text and require the agent to process it under a fixed trusted prompt. 4. Add a validation stage before packaging: - Flag phrases that attempt to override prior instructions. - Flag tool invocation requests, credential requests, external URLs, encoded payloads, and commands unrelated to the source methodology. - Require human approval for generated executable scripts or operational instructions. 5. Ensure generated Skills cannot automatically inherit scripts or commands from source material. Any executable artifact should require explicit user approval and independent code review. 6. Run the distillation agent with least privilege: disable unnecessary network, shell, credential, and filesystem tools while processing untrusted documents. ]]>
