Back to skill

Security audit

Nccl Optimizer

Security checks for vulnerabilities and agentic risk

Overview

The skill’s GPU benchmarking purpose is clear, but its multi-node mode can turn unvalidated node text into shell commands on the local or MPI environment.

Review carefully before installing. This skill is appropriate only in a controlled GPU cluster environment, run as an unprivileged user, and with trusted `nodes=` values. Avoid passing untrusted host strings, avoid running it as root or in broadly privileged containers, and prefer a version that validates hostnames/IPs and uses subprocess argument arrays instead of shell command strings.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
__init__.py:317
Finding

Arbitrary Command Execution Through Unsanitized MPI Node Input

Content
View full analysis

Vulnerability Details

File Location: __init__.py, lines 52–57, 317–325, 329–334, and 479–490
Vulnerability Type: OS command injection
Risk Level: High

Vulnerable Code

User-controlled node values are extracted without validation:

python
def _parse_nodes(message: str) -> list:
    """Extract node list from 'nodes=10.0.0.1,10.0.0.2' or 'hosts=a,b'."""
    m = re.search(r"(?:nodes?|hosts?)=([^\s]+)", message, re.I)
    if m:
        return [n.strip() for n in m.group(1).split(",") if n.strip()]
    return []

The resulting values are interpolated directly into a command string:

python
def _run_internode_allreduce(binary: str, mpirun: str, nodes: list,
                              gpus_per_node: int, best_env: dict) -> tuple:
    """Run all_reduce_perf across multiple nodes via MPI. Returns (bw, raw_output)."""
    host_str = ",".join(f"{n}:{gpus_per_node}" for n in nodes)
    total_ranks = len(nodes) * gpus_per_node
    env_args = " ".join(f"-x {k}={v}" for k, v in best_env.items())
    cmd = (
        f"{mpirun} -np {total_ranks} -H {host_str} "
        f"-x NCCL_DEBUG=WARN {env_args} "
        f"{binary} -b 8M -e 4G -f 2 -g 1 2>&1"
    )
    out = _run_cmd(cmd, timeout=300)
    return _parse_bandwidth(out), out

The command is then executed through a shell:

python
def _run_cmd(cmd: str, timeout: int = 60) -> str:
    """Run *cmd* via shell, return stdout+stderr stripped; empty string on any failure."""
    try:
        return subprocess.check_output(
            cmd, shell=True, stderr=subprocess.STDOUT, text=True, timeout=timeout
        ).strip()
    except Exception:
        return ""

The vulnerable execution path is reached here:

python
peer_nodes = _parse_nodes(message)
mpirun = _find_mpirun()
binary = _find_binary("all_reduce_perf")

if peer_nodes and binary and mpirun:
    all_nodes = [local_hostname] + [n for n
...[truncated 3030 chars]
Remediation
View remediation

Remediation Suggestions

  1. Replace command strings with explicit argument arrays and disable shell interpretation:

    python
    cmd = [
        mpirun,
        "-np", str(total_ranks),
        "-H", host_str,
        "-x", "NCCL_DEBUG=WARN",
    ]
    
    for key, value in best_env.items():
        cmd.extend(["-x", f"{key}={value}"])
    
    cmd.extend([
        binary,
        "-b", "8M",
        "-e", "4G",
        "-f", "2",
        "-g", "1",
    ])
    
    result = subprocess.run(
        cmd,
        shell=False,
        stdout=subprocess.PIPE,
        stderr=subprocess.STDOUT,
        text=True,
        timeout=300,
        check=False,
    )
    
  2. Strictly validate every node before command construction. Accept only valid IPv4 addresses, IPv6 addresses, or hostnames conforming to an explicitly defined policy. Reject shell metacharacters, empty labels, ports, paths, and MPI-specific syntax unless deliberately supported.

  3. Use Python's ipaddress module for IP addresses and a restrictive, length-bounded hostname validator for DNS names.

  4. Avoid building environment-variable assignments into shell text. Use subprocess argument elements for MPI's -x options or the subprocess env parameter where appropriate.

  5. Refactor _run_cmd() to accept only argument sequences and permanently remove shell=True. If shell execution is genuinely required elsewhere, isolate it in a separate function that never receives user-controlled data.

  6. Add regression tests using inputs containing characters such as ;, |, &, backticks, $(), redirection operators, quotes, newlines, and comment markers. Confirm that malformed nodes are rejected before any subprocess starts.

  7. Run the Skill as a dedicated, unprivileged account or tightly restricted container. Do not run it as root or recommend broad --privileged container access when narrow device mappings and capabilities are sufficient.

Vulnerability Patterns
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (5)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
97% confidence
Finding

Using shell=True in a generic command runner is especially dangerous here because the skill later composes commands from multiple dynamic sources, including parsed nodes= input and discovered environment values, making parameter abuse feasible. Since this skill is designed for system benchmarking on GPU clusters, successful abuse could execute arbitrary commands locally and potentially across peer nodes via MPI/SSH, amplifying impact beyond a single host.

Content

Scanner excerpt · __init__.py (reported line 55)May include surrounding context.

python
def _run_cmd(cmd: str, timeout: int = 60) -> str:
    """Run *cmd* via shell, return stdout+stderr stripped; empty string on any failure."""
    try:
        return subprocess.check_output(
            cmd, shell=True, stderr=subprocess.STDOUT, text=True, timeout=timeout
        ).strip()
    except Exception:

Privileged Container / Container Escape

High
Category
Privilege Escalation
Confidence
80% confidence
Finding

Potential security issue detected. Manual review is recommended.

Content

Scanner excerpt · __init__.py (reported line 537)May include surrounding context.

python
lines.append("- `NCCL_BUFFSIZE=8388608` — larger buffers improve large-message throughput")
    lines.append("- `NCCL_SOCKET_NTHREADS=4` + `NCCL_NSOCKS_PERTHREAD=4` — more threads for TCP mode")
    lines.append("- Multi-node: set `MASTER_ADDR` / `MASTER_PORT`, or use `torchrun --rdzv_backend=c10d`")
    lines.append("- In containers: mount `/dev/infiniband` and run with `--privileged` or `--cap-add IPC_LOCK` for RDMA")

    return "\n".join(lines)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill clearly instructs the agent to invoke shell-based system utilities such as nvidia-smi, ibv_devinfo, mpirun, and benchmark binaries, but it declares no explicit permissions or allowed-tools scope. That mismatch can lead to over-broad tool access at runtime, making it easier for the skill to execute unintended commands or operate with more capability than reviewers and users expect.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
96% confidence
Finding

The helper executes arbitrary shell strings with subprocess.check_output(..., shell=True), and this function is used with command strings that incorporate variable data such as interface names, binary paths, and especially user-influenced node lists for MPI execution. In this skill context, the code is intended to probe and benchmark the local and remote environment, so shell execution is expected, but the lack of argument-list invocation and input validation creates command-injection risk if an attacker can influence those values.

Content

Scanner excerpt · __init__.py (reported line 55)May include surrounding context.

python
def _run_cmd(cmd: str, timeout: int = 60) -> str:
    """Run *cmd* via shell, return stdout+stderr stripped; empty string on any failure."""
    try:
        return subprocess.check_output(
            cmd, shell=True, stderr=subprocess.STDOUT, text=True, timeout=timeout
        ).strip()
    except Exception:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

When nodes= is supplied, the skill automatically launches an inter-node MPI benchmark, which can trigger remote execution and SSH-based fan-out without a prominent safety confirmation. In a cluster-tuning skill this behavior is functionally relevant, but it is still risky because users may not appreciate that invoking the skill can cause commands to run across multiple hosts.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.