Install
openclaw skills install @mitsuha-m/nccl-optimizerDetect the optimal NCCL configuration for distributed GPU training on this machine. Checks GPU topology (NVLink/PCIe), whether RDMA (InfiniBand / RoCE) is available, benchmarks intra-node collective bandwidth and peer-to-peer GPU bandwidth, and optionally benchmarks inter-node bandwidth via MPI. Use when: setting up multi-GPU or multi-node training, diagnosing slow collective communication, or tuning NCCL for a new cluster node.
openclaw skills install @mitsuha-m/nccl-optimizerFinds the best NCCL communication configuration for distributed training with clear separation of intra-node and inter-node bandwidth metrics.
nvidia-smi topo -m to detect NVLink vs PCIe.ibv_devinfo PORT_ACTIVE state for InfiniBand/RoCE.
NCCL_IB_* env-vars.NCCL_SOCKET_IFNAME × NCCL_NET_GDR_LEVEL ×
NCCL_IB_TIMEOUT, runs all_reduce_perf -g <N>, picks best bus bandwidth.p2p_bw for GPU↔GPU pair bandwidth (if available).nodes= passed, runs MPI all_reduce_perf across nodes;
otherwise emits a ready-to-run command.| Tool | Purpose | Install |
|---|---|---|
nvidia-smi | GPU info + topology | NVIDIA driver |
ibv_devinfo | RDMA detection | apt install ibverbs-utils |
all_reduce_perf | Collective benchmark | See below |
p2p_bw | Peer-to-peer benchmark | Same nccl-tests build |
mpirun | Inter-node benchmark | apt install openmpi-bin |
git clone https://github.com/NVIDIA/nccl-tests.git
cd nccl-tests
# For V100 (sm_70), A100 (sm_80), A800 (sm_80), H100 (sm_90):
make -j$(nproc) CUDA_HOME=/usr/local/cuda \
NVCC_GENCODE="-gencode=arch=compute_80,code=sm_80"
export PATH=$PWD/build:$PATH
# Intra-node only
openclaw skill run nccl_optimizer
# Include inter-node benchmark (requires passwordless SSH + MPI)
openclaw skill run nccl_optimizer "nodes=10.0.0.1,10.0.0.2"
| Metric | What it measures |
|---|---|
| All-reduce bus BW (intra) | Collective throughput across local GPUs — relevant for single-node training |
| P2P bandwidth | GPU↔GPU direct copy speed (NVLink ≫ PCIe) |
| All-reduce bus BW (inter) | Collective throughput across nodes — bottleneck for multi-node training |
(N-1)/N × data / time. Compare at same N.