Technical
Explorable diagram: an HPC AI cluster
For: engineers and SREs · architectsPrerequisites: Basic notions of GPUs and distributed training.
An AI training cluster stacks five layers: GPU nodes, an interconnect, parallel storage, a scheduler and a software stack. This diagram shows a version reduced to two nodes of 8 NVIDIA H100 SXM GPUs. Each element opens a card: what it is, what to observe there and with which tool, what threatens it. Observe mode places the collectors and their signals; Threats mode lays the six attack vectors from the reference course on the parts they target.
Click any element of the diagram, or walk it with the keyboard (Tab, then Enter or Space): its card opens below the diagram.
Swipe the diagram sideways to see all of it.
Where the collectors sit
The six threat vectors
- Data load
- Forward pass
- Backward pass
- NCCL AllReduce
- Optimizer
- Waiting at the barrier
Illustrative example: durations built from the order of magnitude in chapter 2 (a step of about 42 ms); the straggler takes twice as long for its backward pass. For more depth: the animated training step and the straggler effect simulator.
Try this
Section titled “Try this”- Open a GPU, then the node that holds it. The GPU exposes its counters through dcgm-exporter; the
hpc_joblabel, written by Slurm prolog and epilog scripts, is what ties those counters to a job, then to a user and an account. - Play a step without a straggler, then with one. With a straggler GPU, every other GPU stops at the barrier before the AllReduce: step time aligns on the slowest, while every GPU looks busy.
- Switch to Threats mode and open vector 3. RDMA traffic bypasses the kernel: neither firewalls nor eBPF see it. The only signals are InfiniBand counters, the same ones operations uses.
- Open the model weights. All six vectors converge on them; the collectors that protect them are those of Observe mode.
Go further
Section titled “Go further”- HPC AI: observability and security reference architecture, the course this diagram is drawn from
- HPC AI cluster architecture: the five layers
- Anatomy of a distributed training step: the five phases and the straggler effect
- The HPC AI observability stack: collection, storage, visualization, eBPF
- HPC AI threat surface: the six vectors and how to detect them
- The animated training step: the ring AllReduce and the barrier in detail
- The straggler effect simulator: efficiency lost from 8 to 1024 GPUs