T01 · Skill Instruction Hijacking
- Location
modeio_skill_audit/skill_safety/prompt_payload.py:44- Finding
Untrusted Repository Instructions Are Embedded Verbatim in Model-Facing Prompts
- Content
View full analysis
Vulnerability Details
File Location:
modeio_skill_audit/skill_safety/scanners/prompt.py:42-57,modeio_skill_audit/skill_safety/finding.py:55, andmodeio_skill_audit/skill_safety/prompt_payload.py:44-47
Vulnerability Type: Prompt injection through an untrusted evidence channel
Risk Level: HighRelevant Code:
python # modeio_skill_audit/skill_safety/scanners/prompt.py:42-57 for pattern, rule_id, severity, confidence, why, fix in PROMPT_OVERRIDE_RULES: if pattern.search(raw_line): add_finding( findings, dedupe, layer_state, layer=LAYER_PROMPT, rule_id=rule_id, category="A", severity=severity, confidence=confidence, file_path=rel_path, line=idx, snippet=raw_line, why=why, fix=fix, tags=["prompt-injection", "hierarchy"], exploitability=0.75, reach=0.65, )python # modeio_skill_audit/skill_safety/finding.py:55 "snippet": truncate_snippet(snippet),python # modeio_skill_audit/skill_safety/prompt_payload.py:44-47 lines.append("SCRIPT_SCAN_JSON") lines.append("```json") lines.append(json.dumps(payload, ensure_ascii=False, indent=2)) lines.append("```")Technical Analysis
Repository content is attacker-controlled input. When a prompt-injection pattern is detected, the complete matching line is retained as the finding's
snippet. The prompt-generation flow subsequently serializes findings into a model-facingSCRIPT_SCAN_JSONblock, particularly when--include-full-findingsis enabled.JSON serialization and Markdown code fences provide presentation boundaries but do not create a dependable instruction boundary for a language model. A downstream reviewer can still interpret imperative text inside the serialized evidence as i ...[truncated 2024 chars]
- Remediation
View remediation
Remediation Suggestions
- Treat every repository-derived value, including snippets, paths, descriptions, and OSINT text, as untrusted data in the prompt contract.
- Add explicit instructions immediately before and after the evidence block stating that embedded text must never be obeyed, executed, or interpreted as reviewer instructions.
- Avoid including full raw findings by default. Prefer normalized rule identifiers, hashes, locations, and safely encoded evidence excerpts.
- Encode untrusted snippets into a representation that minimizes instruction-like interpretation, such as escaped JSON strings accompanied by explicit typed metadata.
- Separate trusted reviewer instructions and untrusted evidence through a structured API or separate message roles rather than concatenating both into one text prompt.
- Validate that the scan file belongs to the specified target repository before generating a prompt.
- Add adversarial tests covering direct override instructions, fake system messages, Markdown fence termination attempts, nested JSON, Unicode obfuscation, and instructions distributed across multiple findings.
- Ensure downstream agents operate with minimal tool and secret access even if prompt injection succeeds.
