Back to skill

Security audit

jev-code-review

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent LLM-based code review skill, but it can send reviewed source code to an external Kimi CLI without clear user consent or data-handling disclosure.

Install only if you are comfortable sending reviewed code, diffs, and possibly secrets to whatever `kimi` CLI is available in the environment. Use it on non-sensitive repositories or add explicit consent, redaction, and a local-only option before use in private codebases.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/jev_review.py:71
Finding

Prompt Injection in LLM-Based Code Review Can Manipulate Security Verdicts

Content
View full analysis

Vulnerability Details

File Location: scripts/jev_review.py:71-72; downstream prompt construction occurs in scripts/primitives.py:81-92, with equivalent patterns at scripts/primitives.py:133-146 and scripts/primitives.py:190-203
Vulnerability Type: Untrusted content embedded directly into evaluator prompts
Risk Level: Medium

Vulnerable Code

scripts/jev_review.py:71-72:

python
code_truncated = code[:6000] if len(code) > 6000 else code
state = f"文件: {filename}\n\n代码内容:\n```\n{code_truncated}\n```"

scripts/primitives.py:81-92:

python
prompt = f"""你是 Jev,一个决策模型。你的任务是回答一个是/否问题,并给出概率。

上下文(代码/变更):

{state[:4000]}

text

问题:{instructions}{criteria_str}

请以 JSON 格式回答,只返回 JSON,不要有其他文字:
```json
{{"answer": 0.0-1.0 之间的概率值}}
text

### Technical Analysis

The Skill accepts source code from a file, diff, Git commit, or standard input. It places that content verbatim inside the same natural-language prompt that supplies the downstream LLM's evaluator instructions. Markdown code fences provide presentation formatting but do not establish a security boundary for an LLM.

Consequently, instructions embedded in reviewed repository content can conflict with or override the intended evaluation request. The resulting model response is parsed as JSON and used directly as the security or correctness score. In `scripts/jev_review.py`, these scores determine whether findings are flagged and whether the final recommendation is `PASS`, `WARN`, `REVIEW`, or `BLOCK`.

No independent static analysis, second isolated evaluator, or validation mechanism verifies that a low-risk score is supported by the reviewed code. This makes the Skill's core review decision vulnerable to indirect prompt injection.

### Attack Path

1. An attacker contributes code, a diff, or commit content containing instructions directed at an LLM evaluator.
2. A user invokes this Skill to review that attacker-controll
...[truncated 1255 chars]
Remediation
View remediation

Remediation Suggestions

  • Treat all reviewed source code, diffs, comments, strings, and commit text as untrusted data rather than instructions.
  • Use provider-supported role separation or structured input fields so immutable evaluator instructions are not concatenated with repository content.
  • Explicitly instruct the evaluator that content inside the review-data field must never alter the evaluation procedure. Do not rely on this instruction alone as the security control.
  • Add an independent deterministic analysis stage for high-confidence classes such as hardcoded secrets, unsafe process execution, path traversal, and injection sinks.
  • Consider using multiple isolated evaluators and require evidence-linked findings before issuing PASS.
  • Validate model output semantically rather than only checking whether it contains parseable numeric JSON.
  • Include adversarial prompt-injection fixtures in tests, covering source comments, string literals, diffs, filenames, and commit messages.
  • Fail safely when the response is malformed, suspiciously uniform, unsupported by evidence, or indicates that review instructions were encountered in the submitted content.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (13)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill advertises operational commands that read files and invoke Python from the shell, but it does not declare any tool scope such as permissions or allowed-tools. This creates an authority mismatch where an agent may infer it is acceptable to access local files or execute commands during review, increasing the chance of unintended file exposure or command execution in environments that auto-wire tool use.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The trigger phrases include very broad requests such as 'code review' and '代码审查', which are common in normal developer conversations. In a skill-routing system, this can cause the skill to activate unexpectedly and influence responses in contexts where the user only wanted generic help, potentially expanding access to review workflows or tool-using behavior without clear intent.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The natural-language strings and documentation in the file are entirely Chinese, including the module description and user-facing help text, with no indication that language choice is optional. Under the policy, forcing a specific language without user opt-in is a locale/language policy concern unless the restriction is clearly justified.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/jev_review.py (reported line 34)May include surrounding context.

python
def get_code_from_commit(commit_ref: str) -> str:
    """从 Git 提交获取变更"""
    try:
        result = subprocess.run(
            ["git", "show", commit_ref],
            capture_output=True,
            text=True,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

The script sends source code content into JevPrimitive for evaluation, and the file gives no indication that this processing may involve an external model or service. In a code-review skill, that can expose proprietary code, secrets, or personal data from diffs/commits to a third party without explicit user awareness or consent.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The module claims to perform structured code review, but when any of the main dimensions are flagged, format_output will access data['question'] even though those entries were built from the question spec using the key 'instructions'. This causes a runtime KeyError during reporting, contradicting the script's documented purpose of producing review output for analyzed code.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The primitive includes outbound model-calling capability via a subprocess, which is a meaningful side effect beyond pure local code review logic. In a review skill, exporting user-provided code or diffs to an external tool without clear necessity and controls can expose sensitive source code or secrets.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
89% confidence
Finding

The code launches an external CLI process to handle model interaction, which expands the skill's trust boundary and creates data-flow and execution risk. Although it uses an argument list rather than a shell string, it still sends supplied review content to an external binary and relies on whatever executable named 'kimi' is present in the environment.

Content

Scanner excerpt · scripts/primitives.py (reported line 26)May include surrounding context.

python
"""默认客户端:使用 kimi CLI"""
        import subprocess
        try:
            result = subprocess.run(
                ["kimi", "chat", "--print", prompt],
                capture_output=True,
                text=True,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The code forwards the provided state/prompt to an external CLI-backed LLM without any user-facing warning, consent flow, or redaction step. Because review inputs often contain proprietary code, credentials, or internal context, silent exfiltration to a third-party model endpoint is a real confidentiality risk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The prompts and instructional text are hardcoded in Chinese, which constrains the skill's interaction language without offering a user choice or documenting a justified locale restriction. This matches the language/locale policy concern for natural-language content embedded in code.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This scoring prompt is written entirely in Chinese and does not provide a mechanism for user opt-in or locale selection. Hardcoding a single language in natural-language instructions can violate organizational language/locale policy when no choice or justification is provided.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The choice-selection prompt also forces Chinese-language interaction and output guidance, with no user-configurable language option. This is a repeated natural-language locale restriction present in code string literals.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The file consistently instructs and demonstrates the skill in Chinese, including the invocation example and review questions, but does not indicate that language selection is user-configurable. This can violate language-choice policy if the skill implicitly forces a specific language without user opt-in or justification.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.