Back to skill

Security audit

MetriLLM

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent local LLM benchmarking skill with an optional disclosed leaderboard upload, though users should be careful with the global npm install and shared hardware/model data.

Install only if you are comfortable adding a global npm CLI, preferably without sudo and after checking the package source. Use the share command only when you intentionally want to publish benchmark results plus CPU, RAM, GPU, model name, and scores to the public leaderboard.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:23
Finding

Unpinned Global npm Package Installation Creates Supply-Chain Exposure

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 8 and 23-26
Vulnerability Type: Unpinned third-party dependency installed globally
Risk Level: Medium

Vulnerable Code

yaml
install: npm install -g metrillm
bash
npm install -g metrillm

Technical Analysis

The Skill instructs users to install the latest available version of the metrillm npm package globally. It does not pin a reviewed version, specify an integrity hash, use a lockfile, or require package provenance verification.

Because npm packages may run lifecycle scripts during installation, a compromised package release, compromised maintainer account, or maliciously replaced package could execute code with the permissions of the user running npm. The global installation also places the metrillm executable in a shared command path, increasing the scope and duration of a compromised installation.

The package name is consistent throughout the document, so there is no direct evidence of typosquatting or intentional malicious behavior. The vulnerability is the unsafe, unpinned dependency acquisition process.

Attack Path

  1. An attacker compromises the referenced npm package, its publisher account, or a future package release.
  2. A user follows the Skill instructions and runs npm install -g metrillm.
  3. npm retrieves the current package version without enforcing a reviewed version or integrity value.
  4. Malicious package lifecycle code can execute during installation with the invoking user's privileges.
  5. A malicious global metrillm executable can remain available and execute again during subsequent benchmark commands.

Impact Assessment

A compromised dependency could read or modify files accessible to the invoking user, access environment variables and user-level credentials, execute arbitrary processes, and install a malicious global CLI. If the installation is run with elevated privileges, the impact c ...[truncated 197 chars]

Remediation
View remediation

Remediation Suggestions

  • Pin the dependency to a specific reviewed version, for example:
    bash
    npm install -g metrillm@<reviewed-version>
    
  • Document the expected package publisher, registry, version, and integrity or provenance information.
  • Verify npm package provenance and published checksums before installation.
  • Prefer a project-local dependency with a committed lockfile instead of a global installation.
  • Disable lifecycle scripts when they are not required:
    bash
    npm install --ignore-scripts metrillm@<reviewed-version>
    
  • Explicitly warn users not to run the installation with sudo or an administrator account.
  • Periodically review the pinned package version before updating it.

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:35
Finding

Unquoted Model Argument Permits Shell Word Splitting and Option Injection

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 35, 48, and 61
Vulnerability Type: Unsafe handling of user-controlled command arguments
Risk Level: Medium

Vulnerable Code

bash
metrillm bench --model $ARGUMENTS --json
bash
metrillm bench --model $ARGUMENTS --perf-only --json
bash
metrillm bench --model $ARGUMENTS --share

Technical Analysis

The user-controlled $ARGUMENTS value is expanded without quotation or validation. In a normal Bash expansion, this permits word splitting and pathname expansion. Consequently, a value intended to be one model name can become multiple command-line arguments, including additional options interpreted by the metrillm CLI.

This construction does not, by variable expansion alone in ordinary Bash, cause shell metacharacters contained in the variable to be reparsed as command separators. However, arbitrary command execution could become possible if an Agent framework first performs textual substitution into the command and then submits the resulting string for shell parsing, or if another execution layer uses eval. The directly established risk from the shown code is argument and option injection.

This is particularly relevant because one documented option, --share, uploads benchmark and hardware information. Depending on CLI argument precedence and parsing behavior, injected options could change benchmark behavior or enable functionality that the user did not intend.

Attack Path

  1. An attacker influences the model-name argument supplied to the Skill.
  2. The Skill inserts that value into one of the documented commands as unquoted $ARGUMENTS.
  3. Bash splits the value into multiple words and expands matching pathname patterns.
  4. Additional words beginning with hyphens may be interpreted as separate MetriLLM options rather than as part of the model name.
  5. The benchmark can run with attacker-influenced behavior, poten ...[truncated 1076 chars]
Remediation
View remediation

Remediation Suggestions

  • Quote the model argument in every command:
    bash
    metrillm bench --model "$ARGUMENTS" --json
    metrillm bench --model "$ARGUMENTS" --perf-only --json
    metrillm bench --model "$ARGUMENTS" --share
    
  • Validate the model name against the syntax supported by Ollama or LM Studio before invoking the CLI.
  • Reject control characters, newlines, and unexpected leading option prefixes.
  • Use an argument-array execution API rather than constructing a shell command string.
  • Do not use eval, bash -c with concatenated input, or pre-shell textual interpolation.
  • Require explicit user confirmation immediately before any --share operation and display the categories of hardware and benchmark data that will be transmitted.
  • Where supported by the CLI, use an end-of-options delimiter to prevent model values from being interpreted as options.
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (2)

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill is presented as a local benchmarking tool, but it also documents a --share mode that uploads benchmark data, including hardware specifications, to a public leaderboard. This creates a transparency and data-exposure issue: users may invoke a skill intended for local evaluation without appreciating that it supports external transmission of environment details.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

A public leaderboard upload capability is broader than the core stated purpose of local model evaluation and can expose system metadata such as CPU, RAM, GPU, and model usage externally. Even if no personal data is intended to be sent, hardware fingerprints and model selections may still create privacy, policy, or enterprise data-handling concerns, especially in managed environments.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.