Technical
Simulator: the straggler effect in distributed training
For: engineers and SREs · architectsPrerequisites: Basic notions of distributed training on GPUs.
In synchronous distributed training (data parallelism), each GPU computes gradients on its own batch, then all of them meet at a barrier for the AllReduce that averages them. The step can only finish once the slowest has finished: a single degraded GPU, a straggler, sets the pace for everyone else. And while they wait for it, the GPUs are reported as “busy”.
Prices change fast and vary by contract: this guide gives none; use your provider’s dated price list. Without a price, the simulator expresses the loss in GPU-hours; enter your GPU-hour price, or your internal cost, to see it in euros.
Parameters
Result
- Loading
- Compute
- Waiting at barrier
- AllReduce
Values for the GPUs shown
| GPU | Speed | Compute | Wait |
|---|
Simplified data-parallel model: each GPU computes its batch, all of them wait for the slowest at the barrier, then AllReduce takes the same time for everyone. The ideal assumes every GPU at nominal speed with no variability. The optimizer step and compute/communication overlap are ignored. Variability is a single reproducible draw (fixed seed), not an average: expected efficiency can differ by a few points. Orders of magnitude.
Things to try
Section titled “Things to try”- A single GPU at 70% out of 64 GPUs (the default): efficiency drops to about 76%, that is 367 GPU-hours lost per day out of 1,536 available, and 11,001 over 30 days. Switch to 1,024 GPUs: efficiency is the same, but the GPU-hours lost are multiplied by sixteen (5,867 per day, 176,008 over 30 days). The cost of a straggler grows with cluster size.
- No faulty GPU, just variability: set 0 slowed-down GPUs and 5% variability. Compare 8 GPUs and 1,024 GPUs: displayed efficiency goes from about 99% to about 90%, because the more GPUs, the further the maximum drifts from the mean. The simulator makes a single, reproducible draw (fixed seed), and that draw is a lucky one at 8 GPUs: averaged over many draws, expected efficiency is closer to 95% at 8 GPUs and 89% at 1,024 GPUs. The gap narrows, the trend holds. Even a healthy fleet loses efficiency as it grows, simply because there is always one GPU running a little late.
- Reported utilization versus useful compute: with a straggler, reported utilization stays close to 100% while useful compute drops. No alert based on GPU utilization will ever fire.
- Heavier communication: raise AllReduce to 200 ms. Efficiency versus ideal even improves slightly (the straggler weighs less in a longer step), but useful compute drops further: the network becomes the topic, and optimization is no longer about the GPUs.
The rule to remember
Section titled “The rule to remember”The duration of a synchronous step is the maximum of the compute times, not their mean. The maximum of N random values grows with N: the phenomenon is structural and gets worse with scale. What you need to observe is step time and its breakdown per rank, and the gap between the slowest rank and the median, rather than the GPU utilization rate.
Going further: anatomy of a training step, the technical dimension, the cost of observability for the business and the bridges between dimensions.
Revised on 2 October 2026: prices removed; the loss shows in GPU-hours, and in euros only with your price.