Skip to content

TechnicalPractitioner

Explorable diagram: an HPC AI cluster

For: engineers and SREs · architectsPrerequisites: Basic notions of GPUs and distributed training.

Reading mode

An AI training cluster stacks five layers: GPU nodes, an interconnect, parallel storage, a scheduler and a software stack. This diagram shows a version reduced to two nodes of 8 NVIDIA H100 SXM GPUs. Each element opens a card: what it is, what to observe there and with which tool, what threatens it. Observe mode places the collectors and their signals; Threats mode lays the six attack vectors from the reference course on the parts they target.

Click any element of the diagram, or walk it with the keyboard (Tab, then Enter or Space): its card opens below the diagram.

Swipe the diagram sideways to see all of it.

HPC AI cluster: Slurm scheduler, two nodes of 8 H100 GPUs linked by NVLink, InfiniBand NDR interconnect, parallel storage, model weights and observability planeObservability planecollect · store · GrafanaL4 · Slurm schedulerslurm-exporterGPU node 1 · 8 × H100 SXMdcgm-exporterFalco · TetragonL5 · CUDA · PyTorch · NCCLL5 · PyTorch profiler · NCCL logsGPU 0GPU 1GPU 2GPU 3GPU 4GPU 5GPU 6GPU 7NVLink · 900 GB/sGPU node 2 · 8 × H100 SXMdcgm-exporterFalco · TetragonL5 · CUDA · PyTorch · NCCLL5 · PyTorch profiler · NCCL logsGPU 8GPU 9GPU 10GPU 11GPU 12GPU 13GPU 14GPU 15NVLink · 900 GB/sL2 · InfiniBand NDR · 400 Gb/s · fat tree · RDMAibstat · perfqueryL3 · Parallel storageLustre / IBM Storage ScaleLustre metricsModel weightsasset to protecttargeted by all 6 vectorseBPF · Falco165234

Where the collectors sit

The six threat vectors

  1. Data load
  2. Forward pass
  3. Backward pass
  4. NCCL AllReduce
  5. Optimizer
  6. Waiting at the barrier

Illustrative example: durations built from the order of magnitude in chapter 2 (a step of about 42 ms); the straggler takes twice as long for its backward pass. For more depth: the animated training step and the straggler effect simulator.

  1. Open a GPU, then the node that holds it. The GPU exposes its counters through dcgm-exporter; the hpc_job label, written by Slurm prolog and epilog scripts, is what ties those counters to a job, then to a user and an account.
  2. Play a step without a straggler, then with one. With a straggler GPU, every other GPU stops at the barrier before the AllReduce: step time aligns on the slowest, while every GPU looks busy.
  3. Switch to Threats mode and open vector 3. RDMA traffic bypasses the kernel: neither firewalls nor eBPF see it. The only signals are InfiniBand counters, the same ones operations uses.
  4. Open the model weights. All six vectors converge on them; the collectors that protect them are those of Observe mode.