Back to skill

Security audit

Ollama Model Pilot

Security checks for vulnerabilities and agentic risk

Overview

ModelPilot is a local-only Ollama benchmarking skill with disclosed scripts and safety boundaries; the noted issues are limitations to use carefully, not evidence of malicious behavior.

Install only if you want local Ollama benchmark assistance. Use fictional or explicitly chosen files for prompts, keep benchmark outputs local, review reports manually before replacing a workflow model, and do not rely on reports generated from untrusted or edited JSON files.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/modelpilot_report.py:12
Finding

Incomplete Benchmark Data Can Produce a False Replacement-Ready Decision

Content
View full analysis

Vulnerability Details

File Location: scripts/modelpilot_report.py, lines 12–26
Vulnerability Type: Insufficient integrity and completeness validation
Risk Level: Medium

Vulnerable Code

python
def decision_for_model(records: list[dict[str, Any]], rounds_required: int = 2) -> tuple[str, str]:
    rounds = {record.get("round") for record in records}
    if len(rounds) < rounds_required:
        return "candidate_only", "Only one benchmark round is complete."

    failures = [record for record in records if not record.get("success")]
    format_failures = [record for record in records if not record.get("format_pass")]
    think_leaks = [record for record in records if record.get("think_leak")]

    if failures:
        return "not_recommended", f"{len(failures)} prompt runs failed."
    if think_leaks:
        return "not_recommended", f"{len(think_leaks)} outputs show possible thinking leakage."
    if format_failures:
        return "observe", f"{len(format_failures)} outputs failed the expected format check."
    return "replace_ready", "Two rounds passed mechanical checks. Human semantic review is still required."

Technical Analysis

The replacement decision treats the presence of any two distinct round values as proof that two complete benchmark rounds occurred. It does not validate:

  • That the round identifiers are the expected values, such as rounds 1 and 2.
  • That every configured prompt was executed in each round.
  • That each round contains the same prompt identifiers.
  • That records are unique rather than duplicated.
  • That record coverage agrees with prompt_count and rounds_requested.
  • That model, prompt, round, and result fields have valid types and values.

Consequently, as few as two successful records with different round labels can cause the function to return replace_ready. This contradicts the Skill's documented requirement to run the same fixed prompt set in two complete, independent rounds.

Because ` ...[truncated 1486 chars]

Remediation
View remediation

Remediation Suggestions

  1. Define and enforce a strict schema for benchmark input, including field types and allowed values.
  2. Require exact expected round identifiers, such as every integer from 1 through rounds_required.
  3. Determine the expected prompt-ID set and verify that every model has exactly one record for every prompt in every required round.
  4. Reject duplicate (model, round, prompt_id) records.
  5. Validate record coverage against trusted prompt_count and rounds_requested values.
  6. Reject missing, null, non-Boolean, or incorrectly typed result fields instead of interpreting them through Python truthiness.
  7. Fail closed with an invalid_results or candidate_only decision whenever input completeness cannot be established.
  8. Add tests covering truncated rounds, duplicate records, unexpected round values, inconsistent prompt sets, missing fields, and manipulated metadata.

A hardened decision routine should only return replace_ready after proving that all required prompts completed successfully in every required round.

Vulnerability Patterns
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (3)

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · scripts/modelpilot_benchmark.py (reported line 34)May include surrounding context.

python
for item in prompts:
        if not isinstance(item, dict) or not item.get("id") or not item.get("prompt"):
            raise ValueError("Each prompt must contain at least 'id' and 'prompt'.")
    return prompts


def has_think_leak(text: str) -> bool:

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
85% confidence
Finding

The skill references executable local scripts and describes shell/file-based benchmark workflows, but it does not declare any explicit tool scope such as allowed-tools or permissions. That creates ambiguity about what filesystem and shell capabilities the agent may use, increasing the risk of overbroad local access or unintended command execution when the skill is invoked.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/modelpilot_benchmark.py (reported line 62)May include surrounding context.

python
def run_ollama(model: str, prompt: str, timeout: int) -> tuple[bool, str, str, float]:
    started = time.monotonic()
    try:
        completed = subprocess.run(
            ["ollama", "run", model],
            input=prompt,
            text=True,

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
tests/test_modelpilot_report.py:19