Back to skill

Security audit

Vllm

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent vLLM deployment guide, but its production-style commands use mutable software and expose an unauthenticated GPU-backed inference service without adequate safety scoping.

Review before installing or following the commands. Prefer pinned vLLM package versions and pinned container digests, bind services to localhost or a private interface unless intentionally public, require API authentication, and avoid mounting a Hugging Face cache containing unrelated tokens or private model artifacts.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:26
Finding

Unpinned Third-Party Packages and Mutable Container Image

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 26-32
Vulnerability Type: Unpinned executable dependencies
Risk Level: Medium

Complete Code Snippet:

bash
pip install vllm  # 需要 CUDA 12.1+

# Docker 部署(推荐生产环境)
docker run --runtime nvidia --gpus all \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    -p 8000:8000 vllm/vllm-openai:latest \
    --model meta-llama/Llama-3.1-8B-Instruct

Technical Analysis

The installation instructions execute third-party artifacts without pinning them to immutable, audited versions. pip install vllm allows the package manager to select the current compatible vLLM release and its transitive dependencies. The Docker instruction similarly uses the mutable latest tag rather than an immutable image digest.

Consequently, the effective code executed by these commands may change after the Skill has been reviewed. If a package maintainer account, package registry, container registry, release pipeline, or transitive dependency is compromised, following the documented commands could install and execute attacker-controlled code.

The container is additionally granted access to all NVIDIA GPUs and write access to the host's Hugging Face cache through the bind mount. These permissions increase the impact of a compromised image.

Attack Path

  1. An attacker compromises a relevant package, transitive dependency, registry account, container image, or release pipeline.
  2. The attacker publishes a malicious release that satisfies the unpinned pip request or replaces the image referenced by the mutable latest tag.
  3. A user follows the instructions in SKILL.md.
  4. The package installer or container runtime retrieves the altered artifact.
  5. Attacker-controlled code runs with the invoking user's package-installation privileges or inside a GPU-enabled container.
  6. In the Docker scenario, the malicious process can consume GPU resources and read or ...[truncated 736 chars]
Remediation
View remediation

Remediation Suggestions

  • Pin vLLM to a reviewed, exact version, such as vllm==<approved-version>.
  • Use a lock file or constraints file with cryptographic hashes for vLLM and all transitive Python dependencies.
  • Install only from explicitly trusted package indexes and verify release provenance where available.
  • Replace vllm/vllm-openai:latest with an approved version tag and an immutable digest, for example vllm/vllm-openai:<version>@sha256:<verified-digest>.
  • Scan container images and Python dependencies for known vulnerabilities before deployment.
  • Run the container as a non-root user where supported.
  • Mount the Hugging Face cache read-only unless writes are operationally necessary, or use a dedicated cache containing no unrelated credentials or sensitive artifacts.
  • Limit GPU and filesystem access to the minimum required for the selected model.

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:29
Finding

Inference API Exposed Without Authentication or Network Restrictions

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 29-32
Vulnerability Type: Unauthenticated network service exposure
Risk Level: Medium

Complete Code Snippet:

bash
docker run --runtime nvidia --gpus all \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    -p 8000:8000 vllm/vllm-openai:latest \
    --model meta-llama/Llama-3.1-8B-Instruct

Technical Analysis

The Docker example publishes container port 8000 through -p 8000:8000. Without a host IP restriction, Docker normally publishes the port on all host interfaces. The command does not configure an API key, TLS, an authenticated reverse proxy, a firewall allowlist, rate limiting, or request-size controls.

If the host is reachable from an untrusted network and external firewall controls do not block the port, remote clients may directly access the OpenAI-compatible inference endpoint. Exploitation does not require command injection or compromise of the underlying model server; an attacker can abuse the API using ordinary inference requests.

Attack Path

  1. An administrator launches the documented Docker command on a host reachable by other users or the public Internet.
  2. Docker publishes TCP port 8000 on the host's network interfaces.
  3. No external firewall, reverse proxy, or authentication layer prevents untrusted access.
  4. An attacker discovers the exposed port through network scanning or service enumeration.
  5. The attacker submits unauthenticated inference requests, including large or concurrent workloads.
  6. The requests consume model context capacity, GPU memory, GPU computation, and server concurrency, potentially degrading or denying service to legitimate clients.

Impact Assessment

An unauthenticated attacker could consume the exposed inference service and its associated GPU resources. The practical scope includes unauthorized model use, operational cost increases, workload interference, and denial o ...[truncated 418 chars]

Remediation
View remediation

Remediation Suggestions

  • Bind the published port to loopback when only local access is required:
    bash
    -p 127.0.0.1:8000:8000
    
  • Require an API key using the supported vLLM authentication option and store the key in a secret manager rather than in the Skill or command history.
  • Place remotely accessible deployments behind an authenticated reverse proxy or API gateway.
  • Terminate TLS with a valid certificate and disable plaintext remote access.
  • Restrict source networks using host and cloud firewalls or private-network controls.
  • Configure per-client rate limits, concurrency limits, request-size limits, timeouts, and maximum token limits.
  • Add monitoring and alerts for unusual request volume, sustained GPU saturation, repeated failures, and unauthorized access attempts.
  • Document that direct public exposure of the vLLM service is not a production-safe default.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (2)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The manifest description and main instruction text are written to position the skill as a Chinese-language assistant, with no indication that users may choose another language. This can violate language/locale policy requirements when a skill implicitly enforces one language without explicit opt-in or documented justification.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
92% confidence
Finding

The Docker example uses a mutable image reference (vllm/vllm-openai:latest), which can change over time and cause users to pull unreviewed or compromised content. In an infrastructure/deployment skill, this is more dangerous because readers are likely to copy-paste the command directly into production or GPU hosts with broad system access.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.