Skip to content

TechnicalPractitioner

Animated diagram: a distributed training step

For: engineers and SREs · architectsPrerequisites: Basic notions of training models on GPUs.

Reading mode

In data-parallel distributed training, each GPU computes gradients on its share of the batch, then all GPUs average them before updating the model. That average, the AllReduce, imposes a barrier: nobody can move on until the slowest one is done. The animation plays one step on 8 GPUs, phase by phase, and shows what the usual dashboards do not. The phases are detailed in the chapter Anatomy of a distributed training step.

Parameters

loadingforwardbackwardwaitingAllReduceoptimizer

The 8 GPUs and the ring AllReduce

Step timeline, GPU by GPU

What the step costs

Step duration
ideal without straggler
Scaling efficiency
ideal duration / actual duration
Waiting at the barrier
total over the 8 GPUs, per step
AllReduce share
of step time
Reported GPU utilization
GPU "busy", including while waiting and communicating (DCGM_FI_DEV_GPU_UTIL, nvidia-smi)
Useful compute
nominal compute × active SM share, divided by step time (close to DCGM_FI_PROF_SM_ACTIVE)

Illustrative durations: loading 6 ms, forward 10 ms, backward 16 ms, AllReduce 10 ms, optimizer 2 ms; the slowdown applies to the straggler's forward and backward passes. Simplification: phases are sequential; in practice PyTorch DDP overlaps part of the AllReduce with the backward pass (gradient buckets), which shrinks the wait without removing it.

  1. Watch a step without a straggler. During the AllReduce, each GPU sends one eighth of its gradients to its neighbor, fourteen times in a row: seven reduce-scatter steps (each GPU’s small cells fill up as contributions add up), then seven all-gather steps (each GPU receives the already averaged chunks). Use “Next step” to follow the exchanges one at a time.
  2. Slow GPU 2 down to 1.5×. The other seven finish their backward pass and wait, hatched, at the barrier. Step time goes from 44 to 57 ms and efficiency drops to about 77%. Yet reported GPU utilization goes up: waiting counts as busy.
  3. Compare the two gauges. At 2×, reported utilization exceeds 90% while useful compute falls below 25%. That is the gap in the illustrative example of the reference chapter: 87% reported for about 19% of real compute.
  4. Turn on prefetching. Data loading overlaps with compute and almost vanishes from the timeline: the step gets shorter and useful compute improves. With a 2× straggler, the gain is still 5 ms while the straggler costs 26: you do not fix a synchronization problem with an I/O optimization.

Durations are orders of magnitude chosen for readability, and phases are strictly sequential. In practice, PyTorch DDP launches the AllReduce in gradient buckets during the backward pass, hiding part of the communication, and NCCL splits each send into finer chunks. The principle holds: a ring AllReduce moves about 2 × (N - 1) / N times the gradient size per GPU, and the barrier aligns everyone on the slowest.

QuestionWhat to measureWith which tool
Is the step slowing down?step time, per rank and mediantraining loop instrumentation (timestamp per step and per rank), exported to a time series database (VictoriaMetrics, Prometheus or a SaaS equivalent)
Is there a straggler?gap between the slowest rank and the mediansame metric, alert on the gap rather than the average
Is the GPU really computing?SM activity, occupancy, tensor core usageDCGM profiling metrics: DCGM_FI_PROF_SM_ACTIVE, DCGM_FI_PROF_SM_OCCUPANCY, DCGM_FI_PROF_PIPE_TENSOR_ACTIVE, through dcgm-exporter
Where does step time go?breakdown into loading, compute, communicationPyTorch Profiler or Nsight Systems on a few steps, not continuously
Is the network to blame?collective latency, retransmissions and port waits on InfiniBandNCCL logs (NCCL_DEBUG=INFO), InfiniBand port counters

DCGM_FI_DEV_GPU_UTIL, the equivalent of the utilization column in nvidia-smi, only measures the share of time a kernel is running: useful to spot an idle GPU, misleading for judging efficiency. The full stack is described in The HPC AI observability stack, and the straggler effect simulator puts a cost on it at cluster scale.