T09 · Insecure Skill Coding Practices
- Location
scripts/jev_review.py:71- Finding
Prompt Injection in LLM-Based Code Review Can Manipulate Security Verdicts
- Content
View full analysis
Vulnerability Details
File Location:
scripts/jev_review.py:71-72; downstream prompt construction occurs inscripts/primitives.py:81-92, with equivalent patterns atscripts/primitives.py:133-146andscripts/primitives.py:190-203
Vulnerability Type: Untrusted content embedded directly into evaluator prompts
Risk Level: MediumVulnerable Code
scripts/jev_review.py:71-72:python code_truncated = code[:6000] if len(code) > 6000 else code state = f"文件: {filename}\n\n代码内容:\n```\n{code_truncated}\n```"scripts/primitives.py:81-92:python prompt = f"""你是 Jev,一个决策模型。你的任务是回答一个是/否问题,并给出概率。 上下文(代码/变更):{state[:4000]}
text 问题:{instructions}{criteria_str} 请以 JSON 格式回答,只返回 JSON,不要有其他文字: ```json {{"answer": 0.0-1.0 之间的概率值}}text ### Technical Analysis The Skill accepts source code from a file, diff, Git commit, or standard input. It places that content verbatim inside the same natural-language prompt that supplies the downstream LLM's evaluator instructions. Markdown code fences provide presentation formatting but do not establish a security boundary for an LLM. Consequently, instructions embedded in reviewed repository content can conflict with or override the intended evaluation request. The resulting model response is parsed as JSON and used directly as the security or correctness score. In `scripts/jev_review.py`, these scores determine whether findings are flagged and whether the final recommendation is `PASS`, `WARN`, `REVIEW`, or `BLOCK`. No independent static analysis, second isolated evaluator, or validation mechanism verifies that a low-risk score is supported by the reviewed code. This makes the Skill's core review decision vulnerable to indirect prompt injection. ### Attack Path 1. An attacker contributes code, a diff, or commit content containing instructions directed at an LLM evaluator. 2. A user invokes this Skill to review that attacker-controll ...[truncated 1255 chars]- Remediation
View remediation
Remediation Suggestions
- Treat all reviewed source code, diffs, comments, strings, and commit text as untrusted data rather than instructions.
- Use provider-supported role separation or structured input fields so immutable evaluator instructions are not concatenated with repository content.
- Explicitly instruct the evaluator that content inside the review-data field must never alter the evaluation procedure. Do not rely on this instruction alone as the security control.
- Add an independent deterministic analysis stage for high-confidence classes such as hardcoded secrets, unsafe process execution, path traversal, and injection sinks.
- Consider using multiple isolated evaluators and require evidence-linked findings before issuing
PASS. - Validate model output semantically rather than only checking whether it contains parseable numeric JSON.
- Include adversarial prompt-injection fixtures in tests, covering source comments, string literals, diffs, filenames, and commit messages.
- Fail safely when the response is malformed, suspiciously uniform, unsupported by evidence, or indicates that review instructions were encountered in the submitted content.
