Skip to content

TechnicalPractitioner

Anatomy of a distributed training step

For: engineers and SREs · architectsPrerequisites: Have read chapter 1 of the course.

PhaseWhat happensIf it is slow
1. Data load and host to deviceworkers read from Lustre, tokenize, load into pinned CPU memory, then DMA transfer to the GPU over PCIe (roughly 16 to 64 GB/s depending on generation); must overlap with the previous batchthe workload is I/O bound
2. Forward passdata goes through the model layer by layer; each layer is mostly a matrix multiplication, tiled across Streaming Multiprocessors; activations stay in HBM for the backward passcompute or memory saturated
3. Backward passloss computation, then backpropagation layer by layer; each GPU computes gradients on its own data subsetsame
4. AllReduce (NCCL)a ring AllReduce averages gradients across all GPUs over InfiniBand; its share of step time varies widely with model size, network and overlap with computenetwork or straggler
5. Optimizer stepAdamW updates all parameters in parallel; the cycle restarts on the next batchrarely the bottleneck
0 ms ~8 ms ~16 ms ~28 ms ~42 ms
|-----------|--------------|-------------------|-----------------------|
GPU 0 [data load ][ forward pass ][ backward pass ][ NCCL AllReduce ][opt]
GPU 1 [data load ][ forward pass ][ backward pass ][ NCCL AllReduce ][opt]
GPU 2 [data load ][ forward pass ][ backward pass ........ ][ AllReduce ][opt] <- straggler
GPU 3 [data load ][ forward pass ][ backward pass ][ waiting... ][opt]
... ^ synchronization barrier

Every GPU waits at the AllReduce barrier for the slowest one to finish. A single slow GPU is enough to hold back all 32.

Why 87% utilization does not mean 87% compute

Section titled “Why 87% utilization does not mean 87% compute”

Illustrative example, with values constructed for the explanation (not a measurement):

IndicatorValue
Reported GPU utilization87%
Time actually spent computingabout 19% of the step
Time spent in NCCL AllReduce43% of the step

GPU utilization measures whether the GPU has something running, not whether it is computing usefully. A GPU waiting for a collective to finish counts as busy. A straggler can thus cut a large share of global efficiency without triggering a single alert; the simulator below lets you put figures on scenarios.

Revised on 2 October 2026: the AllReduce share of step time is described as variable instead of an unsourced range, the 87% example is presented as constructed for the explanation rather than taken from profiling, and straggler efficiency loss points to the simulator instead of an unsourced percentage.