Technical
Anatomy of a distributed training step
For: engineers and SREs · architectsPrerequisites: Have read chapter 1 of the course.
The five phases of a step
Section titled “The five phases of a step”| Phase | What happens | If it is slow |
|---|---|---|
| 1. Data load and host to device | workers read from Lustre, tokenize, load into pinned CPU memory, then DMA transfer to the GPU over PCIe (roughly 16 to 64 GB/s depending on generation); must overlap with the previous batch | the workload is I/O bound |
| 2. Forward pass | data goes through the model layer by layer; each layer is mostly a matrix multiplication, tiled across Streaming Multiprocessors; activations stay in HBM for the backward pass | compute or memory saturated |
| 3. Backward pass | loss computation, then backpropagation layer by layer; each GPU computes gradients on its own data subset | same |
| 4. AllReduce (NCCL) | a ring AllReduce averages gradients across all GPUs over InfiniBand; its share of step time varies widely with model size, network and overlap with compute | network or straggler |
| 5. Optimizer step | AdamW updates all parameters in parallel; the cycle restarts on the next batch | rarely the bottleneck |
One step across four nodes and 32 GPUs
Section titled “One step across four nodes and 32 GPUs” 0 ms ~8 ms ~16 ms ~28 ms ~42 ms |-----------|--------------|-------------------|-----------------------| GPU 0 [data load ][ forward pass ][ backward pass ][ NCCL AllReduce ][opt] GPU 1 [data load ][ forward pass ][ backward pass ][ NCCL AllReduce ][opt] GPU 2 [data load ][ forward pass ][ backward pass ........ ][ AllReduce ][opt] <- straggler GPU 3 [data load ][ forward pass ][ backward pass ][ waiting... ][opt] ... ^ synchronization barrierEvery GPU waits at the AllReduce barrier for the slowest one to finish. A single slow GPU is enough to hold back all 32.
Why 87% utilization does not mean 87% compute
Section titled “Why 87% utilization does not mean 87% compute”Illustrative example, with values constructed for the explanation (not a measurement):
| Indicator | Value |
|---|---|
| Reported GPU utilization | 87% |
| Time actually spent computing | about 19% of the step |
| Time spent in NCCL AllReduce | 43% of the step |
GPU utilization measures whether the GPU has something running, not whether it is computing usefully. A GPU waiting for a collective to finish counts as busy. A straggler can thus cut a large share of global efficiency without triggering a single alert; the simulator below lets you put figures on scenarios.
Revised on 2 October 2026: the AllReduce share of step time is described as variable instead of an unsourced range, the 87% example is presented as constructed for the explanation rather than taken from profiling, and straggler efficiency loss points to the simulator instead of an unsourced percentage.