T09 · Insecure Skill Coding Practices
- Location
src/input_safety_guard/pipeline.py:294- Finding
Untrusted user input is directly interpolated into the stage-2 classifier prompt
- Content
View full analysis
User: {{user_input}} [/INST] ``` ```python def build_stage2_prompt(self, text: str) -> str: overlay = STAGE2_PROFILE_OVERLAYS[self.stage2_profile].strip() prompt = STAGE2_PROMPT_TEMPLATE.replace("{{user_input}}", text) return prompt.replace("", f"\n{overlay}\n\n\n") ``` ### Technical Analysis The pipeline inserts raw, attacker-controlled text directly into a trusted instruction-formatted prompt. The value is not escaped, structurally separated using model message roles, or encoded as inert data. An attacker can therefore include prompt delimiters or instruction-like text such as ``, `[/INST]`, or semantically equivalent instructions intended to alter the classifier's behavior. The deterministic prefilter reduces exposure but does not eliminate this risk because its keyword and regular-expression rules cannot reliably recognize every paraphrased, obfuscated, multilingual, encoded, or delimiter-based injection. Allowlisted requests also continue to stage 2, so the allowlist does not establish a security boundary. The prompt-injection wording at `pipeline.py:59` is a defensive category description and is appropriate for the Skill's declared purpose. The vulnerability arises from placing untrusted content into the same prompt channel as those trusted classifier instructions. ### Attack Path 1. An attacker prepares an unsafe request that avoids the deterministic stage-1 keywords and regular expressions. 2. The request includes delimiter-like content or indirect instructions intended to make the classifier return a safe result. 3. Stage 1 returns `allow` or `re ...[truncated 1008 chars]- Remediation
View remediation
` as a security boundary. 3. If the classifier only accepts one string: - Serialize the input as a JSON value or another unambiguous data representation. - Escape or reject reserved prompt delimiters. - Clearly instruct the classifier to treat the serialized value exclusively as data. - Avoid decoding and reinserting it into an instruction-bearing section before classification. 4. Use constrained or schema-based classifier output where supported. 5. Add adversarial tests covering: - `[/INST]` and conversation-closing delimiters. - Unicode and whitespace obfuscation. - Multilingual and indirect prompt injection. - Encoded instructions. - Inputs containing fake classifier output. 6. Retain conservative failure behavior so any parsing or classifier error cannot invoke the downstream responder. ]]>
