Technical
Animated diagram: a distributed training step
For: engineers and SREs · architectsPrerequisites: Basic notions of training models on GPUs.
In data-parallel distributed training, each GPU computes gradients on its share of the batch, then all GPUs average them before updating the model. That average, the AllReduce, imposes a barrier: nobody can move on until the slowest one is done. The animation plays one step on 8 GPUs, phase by phase, and shows what the usual dashboards do not. The phases are detailed in the chapter Anatomy of a distributed training step.
Parameters
The 8 GPUs and the ring AllReduce
Step timeline, GPU by GPU
What the step costs
Illustrative durations: loading 6 ms, forward 10 ms, backward 16 ms, AllReduce 10 ms, optimizer 2 ms; the slowdown applies to the straggler's forward and backward passes. Simplification: phases are sequential; in practice PyTorch DDP overlaps part of the AllReduce with the backward pass (gradient buckets), which shrinks the wait without removing it.
Try this
Section titled “Try this”- Watch a step without a straggler. During the AllReduce, each GPU sends one eighth of its gradients to its neighbor, fourteen times in a row: seven reduce-scatter steps (each GPU’s small cells fill up as contributions add up), then seven all-gather steps (each GPU receives the already averaged chunks). Use “Next step” to follow the exchanges one at a time.
- Slow GPU 2 down to 1.5×. The other seven finish their backward pass and wait, hatched, at the barrier. Step time goes from 44 to 57 ms and efficiency drops to about 77%. Yet reported GPU utilization goes up: waiting counts as busy.
- Compare the two gauges. At 2×, reported utilization exceeds 90% while useful compute falls below 25%. That is the gap in the illustrative example of the reference chapter: 87% reported for about 19% of real compute.
- Turn on prefetching. Data loading overlaps with compute and almost vanishes from the timeline: the step gets shorter and useful compute improves. With a 2× straggler, the gain is still 5 ms while the straggler costs 26: you do not fix a synchronization problem with an I/O optimization.
What the animation simplifies
Section titled “What the animation simplifies”Durations are orders of magnitude chosen for readability, and phases are strictly sequential. In practice, PyTorch DDP launches the AllReduce in gradient buckets during the backward pass, hiding part of the communication, and NCCL splits each send into finer chunks. The principle holds: a ring AllReduce moves about 2 × (N - 1) / N times the gradient size per GPU, and the barrier aligns everyone on the slowest.
What to measure for real
Section titled “What to measure for real”| Question | What to measure | With which tool |
|---|---|---|
| Is the step slowing down? | step time, per rank and median | training loop instrumentation (timestamp per step and per rank), exported to a time series database (VictoriaMetrics, Prometheus or a SaaS equivalent) |
| Is there a straggler? | gap between the slowest rank and the median | same metric, alert on the gap rather than the average |
| Is the GPU really computing? | SM activity, occupancy, tensor core usage | DCGM profiling metrics: DCGM_FI_PROF_SM_ACTIVE, DCGM_FI_PROF_SM_OCCUPANCY, DCGM_FI_PROF_PIPE_TENSOR_ACTIVE, through dcgm-exporter |
| Where does step time go? | breakdown into loading, compute, communication | PyTorch Profiler or Nsight Systems on a few steps, not continuously |
| Is the network to blame? | collective latency, retransmissions and port waits on InfiniBand | NCCL logs (NCCL_DEBUG=INFO), InfiniBand port counters |
DCGM_FI_DEV_GPU_UTIL, the equivalent of the utilization column in nvidia-smi, only measures the share of time a kernel is running: useful to spot an idle GPU, misleading for judging efficiency. The full stack is described in The HPC AI observability stack, and the straggler effect simulator puts a cost on it at cluster scale.