Install
openclaw skills install skills-sh:nvidia/skills/nemo-mbridge-perf-expert-parallel-overlapMoE Expert-Parallel Overlap Skill ## References - Stable docs: @docs/training/communication-overlap.md - Structured metadata: @skills/nemo-mbridge-perf-expert-parallel-overlap/card.yaml ## What It Is Expert-parallel (EP) overlap hides the cost of token dispatch/combine…
openclaw skills install skills-sh:nvidia/skills/nemo-mbridge-perf-expert-parallel-overlapExpert-parallel (EP) overlap hides the cost of token dispatch/combine all-to-all
communication by running it concurrently with expert FFN compute. Optionally,
delayed expert weight-gradient computation (delay_wgrad_compute) provides
additional overlap by deferring wgrad to overlap with the next layer's forward.
Bridge supports two dispatcher paths:
| Dispatcher | Backend | When to use |
|---|---|---|
alltoall | Standard MoE all-to-all | Default, broadest compatibility |
flex | DeepEP or HybridEP | Higher overlap on Ampere/Hopper/Blackwell |
Use EP overlap when:
EP > 1Prefer:
alltoall dispatcher for the first rollout (broader compatibility)flex + DeepEP/HybridEP when running on supported GPUs and seeking
additional gainsAvoid EP overlap when:
moe_shared_expert_overlap is enabledExpected outcome:
For the plain EP-overlap isolation benchmark, keep flex dispatch and delayed
wgrad disabled. The measured shape was Qwen3 MoE 30B-A3B SFT on 16 H100 GPUs:
EP=16, alltoall, BF16, global batch size 1024, CUDA graphs disabled,
moe_permute_fusion=false, measured over iterations 3-8.
Use these overrides for the plain-overlap case:
--cuda_graph_impl none \
--moe_flex_dispatcher_backend None \
--moe_a2a_overlap false \
comm_overlap.overlap_moe_expert_parallel_comm=true \
comm_overlap.delay_wgrad_compute=false \
model.moe_shared_expert_overlap=false
Do not use --moe_a2a_overlap true for this isolation test: the performance
harness helper enables both overlap_moe_expert_parallel_comm and
delay_wgrad_compute, so it does not isolate plain EP overlap.
Steady-window timing from that benchmark:
| Case | Steady mean | Relative |
|---|---|---|
| no EP overlap | 41.25s | 1.000x |
| EP overlap | 31.31s | 1.317x |
EP overlap plus delay_wgrad_compute | 31.20s | 1.322x |
This is evidence for enabling plain EP overlap on this inter-node all-to-all shape. It does not show a meaningful independent win from delayed wgrad, and it does not validate fused MoE permutation because that path was disabled for the runtime stack.
A 2026-07-25 controlled Qwen3 30B-A3B pretraining comparison validated plain EP overlap with the production HybridEP path:
Hardware: 16×H100
Precision: BF16
Sequence: 4096
Parallelism: TP1 / PP1 / CP1 / EP16
Batch: MBS1 / GBS1024
Routing: force balance
Dispatcher: flex + HybridEP
CUDA graph: Transformer Engine scopes moe_router + moe_preprocess
Delayed wgrad: disabled
| Case | Steady window | Step time | Model TFLOPS/GPU |
|---|---|---|---|
| overlap off | iterations 5-20 | 24.7138s | 244.039 |
| overlap on, search run | iterations 5-20 | 21.0725s | 286.208 |
| overlap on, independent validation | iterations 41-50 | 20.9920s | 287.305 |
The independent run reduced step time by 15.059% and raised throughput by 17.729% over the reproduced baseline. Loss was finite, skipped and NaN iterations remained zero, and rank-0 peak allocated memory was 62.166 GiB.
A matched Nsight Systems comparison captured the same 463,348 rank-0 kernels per case. Enabling overlap increased communication concurrent with GEMM and attention from 9.079ms (0.11% of communication time) to 3,958.997ms (36.55%). GPU-active interval union fell from 22.821s to 21.221s.
Use this as evidence for the mechanism, not as a universal speedup promise. The dispatcher, graph scopes, routing, parallelism, batch shape, and runtime were held fixed while only plain EP overlap changed.
cfg.comm_overlap.overlap_moe_expert_parallel_comm = True
cfg.comm_overlap.delay_wgrad_compute = False
cfg.model.moe_shared_expert_overlap = False
cfg.model.expert_model_parallel_size = 8
cfg.model.num_moe_experts = 64
cfg.model.moe_token_dispatcher_type = "alltoall"
cfg.model.bf16 = True
cfg.model.fp16 = False
Enable delay_wgrad_compute=True only after the plain overlap path is known to
work and its extra compatibility constraints have been checked.
from megatron.bridge.training.flex_dispatcher_backend import apply_flex_dispatcher_backend
cfg.comm_overlap.overlap_moe_expert_parallel_comm = True
cfg.comm_overlap.delay_wgrad_compute = False
cfg.model.moe_shared_expert_overlap = False
apply_flex_dispatcher_backend(cfg.model, moe_flex_dispatcher_backend="deepep")
# or: apply_flex_dispatcher_backend(cfg.model, moe_flex_dispatcher_backend="hybridep")
Benchmark plain EP overlap first. Enable delay_wgrad_compute=True only as a
separate follow-up A/B after its CUDA-graph and TE compatibility constraints
are satisfied.
expert_model_parallel_size > 1num_moe_experts > 1moe_token_dispatcher_type must be "alltoall" or "flex"moe_shared_expert_overlap = False>= 2.6.0PP > 1, virtual_pipeline_model_parallel_size must be setrecompute_granularity != "full", recompute_method = None,
recompute_num_layers = Nonemtp_num_layers must be None or 1delay_wgrad_compute requires overlap_moe_expert_parallel_comm as a
prerequisitedelay_wgrad_compute with overlap_grad_reduce requires TE >= 2.7.0delay_wgrad_compute with gradient_accumulation_fusion requires TE >= 2.7.0attn scope + delay_wgrad_compute requires TE >= 2.12.0,
gradient_accumulation_fusion = True, and no attention biascfg.comm_overlap.overlap_moe_expert_parallel_comm = True
cfg.comm_overlap.delay_wgrad_compute = False
cfg.model.expert_model_parallel_size = 4
cfg.model.num_moe_experts = 64
cfg.model.moe_token_dispatcher_type = "alltoall"
cfg.model.moe_shared_expert_overlap = False
cfg.model.bf16 = True
Use this as the correctness-first starting point. Add delayed wgrad, flex dispatch, and CUDA-graph interactions only after the plain overlap path is known to work.
Performance harness example inside a Slurm allocation. Keep the model, parallelism, dispatcher, and runtime fixed, and vary only the two overlap overrides:
uv run python scripts/performance/run_script.py \
-m qwen \
-mr qwen3_30b_a3b \
--task pretrain \
-g h100 \
-c bf16 \
-ng 16 \
-gn 8 \
--max_steps 8 \
--cuda_graph_impl none \
--moe_flex_dispatcher_backend None \
--moe_a2a_overlap false \
--tokenizer_type NullTokenizer \
comm_overlap.overlap_moe_expert_parallel_comm=true \
comm_overlap.delay_wgrad_compute=false \
model.moe_shared_expert_overlap=false
Do not use --moe_a2a_overlap true when separating plain EP overlap from
delayed wgrad: the performance harness helper enables both
overlap_moe_expert_parallel_comm and delay_wgrad_compute.
Unit test verification:
uv run python -m pytest \
tests/unit_tests/training/test_comm_overlap.py -k "moe" \
tests/unit_tests/training/test_deepep.py -q
uv run python -m pytest \
tests/unit_tests/training/test_comm_overlap.py \
tests/unit_tests/training/test_deepep.py -q
After a successful run with EP overlap:
CommOverlapConfig finalizationoverlap_moe_expert_parallel_comm appears as True in the logged
configmoe_token_dispatcher_type = "flex" and
the correct backend in logsUse an unprofiled steady window for the throughput acceptance result. Use a matched profile to explain the mechanism:
if self.user_comm_overlap_cfg.overlap_moe_expert_parallel_comm is True:
assert model_cfg.expert_model_parallel_size > 1, ...
assert model_cfg.num_moe_experts > 1, ...
assert model_cfg.moe_token_dispatcher_type in ["alltoall", "flex"], ...
assert model_cfg.bf16 or model_cfg.fp16, ...
assert is_torch_min_version("2.6.0"), ...
# ... PP + VPP check, recompute checks, shared_expert_overlap check ...
if self.user_comm_overlap_cfg.delay_wgrad_compute is True:
# TE version checks for overlap_grad_reduce and gradient_accumulation_fusion
# CUDA graph scope validations for delayed wgrad
assert overlap_moe_expert_parallel_comm, ...
def apply_flex_dispatcher_backend(...):
# GPU architecture check for DeepEP / HybridEP
model_config.moe_token_dispatcher_type = "flex"
model_config.moe_flex_dispatcher_backend = moe_flex_dispatcher_backend
model_config.moe_shared_expert_overlap = False
def _set_moe_a2a_overlap_overrides(recipe, moe_a2a_overlap=False):
if moe_a2a_overlap:
recipe.comm_overlap.overlap_moe_expert_parallel_comm = True
recipe.comm_overlap.delay_wgrad_compute = True
recipe.model.moe_shared_expert_overlap = False
| File | Coverage |
|---|---|
tests/unit_tests/training/test_comm_overlap.py | EP overlap validation, delayed wgrad, CUDA graph + wgrad interaction |
tests/unit_tests/training/test_deepep.py | DeepEP/HybridEP helper activation and GPU gating |
| Symptom | Likely Cause | How To Confirm | Fix |
|---|---|---|---|
assert expert_model_parallel_size > 1 | EP not configured | Check expert_model_parallel_size | Set EP > 1 |
assert moe_token_dispatcher_type | Wrong dispatcher | Check dispatcher type | Use "alltoall" or "flex" |
| assert on BF16/FP16 | Wrong precision | Check bf16 and fp16 | Set bf16 = True |
| hang during training | PyTorch < 2.6 | Check PyTorch version | Upgrade to >= 2.6.0 |
assert virtual_pipeline_model_parallel_size | PP > 1 without VPP | Check PP and VPP config | Set VPP when PP > 1 |
assert recompute_granularity | Full recompute enabled | Check recompute settings | Disable full recompute |
assert overlap_moe_expert_parallel_comm required | delayed wgrad without EP overlap | Check delay_wgrad_compute without overlap | Enable EP overlap first |
assert gradient_accumulation_fusion | CUDA graph + delayed wgrad | Check graph scope + wgrad settings | Enable gradient_accumulation_fusion |
| assert on attention bias | CUDA graph attn + delayed wgrad + bias | Check add_bias_linear / add_qkv_bias | Disable attention bias |
| no throughput gain from flex dispatcher | apply_flex_dispatcher_backend not called | Check moe_token_dispatcher_type in logs | Call apply_flex_dispatcher_backend(...) |
| DeepEP/HybridEP silently skipped | Unsupported GPU | Check warning logs | Run on Ampere/Hopper/Blackwell |
| summed kernel time increases after overlap | Expected concurrency contention or a regression | Compare interval unions, comm/compute intersection, and unprofiled step time | Judge overlap from exposed wall time, not summed per-stream duration |
moe_flex_dispatcher_backend alone does not activate flex dispatch —
you must call apply_flex_dispatcher_backend(...).Last signature refresh: 2026-08-03.
08ea07e0d730