Technical
HPC AI cluster architecture
For: engineers and SREs · architectsPrerequisites: Basic HPC notions: node, scheduler, interconnect.
Physical topology
Section titled “Physical topology”flowchart TB
SCHED["Slurm / PBS scheduler"]
subgraph N1["GPU node 1"]
direction LR
G0["GPU 0<br/>80 GB HBM"] --- G1["GPU 1"] --- G2["GPU 2"] --- G3["GPU 3"]
end
subgraph N2["GPU node 2"]
direction LR
G4["GPU 4"] --- G5["GPU 5"] --- G6["GPU 6"] --- G7["GPU 7"]
end
IB["InfiniBand / RoCE<br/>latency in the order of a microsecond"]
FS["Parallel storage<br/>Lustre / IBM Storage Scale / BeeGFS"]
SCHED --> N1
SCHED --> N2
N1 <--> IB
N2 <--> IB
IB <--> FS
Inside a node, GPUs talk over NVLink; between nodes, over InfiniBand or RoCE. Parallel storage feeds everyone, and the scheduler decides what runs where.
The five layers
Section titled “The five layers”| Layer | What it contains | What to observe | What to protect |
|---|---|---|---|
| L1. GPU nodes | 4 to 8 GPUs per node linked by NVLink (600 GB/s on A100, 900 GB/s on H100, 1.8 TB/s on Blackwell), 80 GB HBM per GPU on A100 and H100, more on later generations (141 GB on H200), 2 orchestration CPUs, hundreds of GB of RAM, local NVMe for staging | dcgm-exporter: SM occupancy, HBM bandwidth, ECC errors, throttling, power | firmware, GPU memory between jobs |
| L2. Interconnect | InfiniBand HDR (200 Gb/s), NDR (400 Gb/s) or XDR (800 Gb/s, Blackwell generation), or RoCE; non-blocking fat-tree topology; gradients flow here during NCCL AllReduce | ibstat, perfquery, port counters, retransmits | the RDMA blind spot |
| L3. Parallel storage | Lustre or IBM Storage Scale (formerly Spectrum Scale, or GPFS), files spread over dozens of servers, aggregate throughput from tens of GB/s to several TB/s depending on system size; GPUs must never wait for data | Lustre throughput, metadata latency | training data integrity |
| L4. Scheduler | Slurm or PBS: queues, GPU allocation, job isolation; prolog and epilog scripts to inject instrumentation | slurm-exporter: jobs, queues, wait time | isolation between jobs and tenants |
| L5. ML software stack | CUDA or ROCm driver, PyTorch or JAX, NCCL for collective communication, Docker or Apptainer (formerly Singularity) containers | PyTorch profiler, NCCL logs | software supply chain |
Four families of signals
Section titled “Four families of signals”| Family | Examples |
|---|---|
| GPU | SM utilization, HBM bandwidth, ECC errors, clock throttling |
| Network | InfiniBand port counters, retransmit rate, per-link throughput |
| Storage | Lustre throughput, metadata latency |
| Model | loss, gradient norm, step time, AllReduce latency |
The fourth family is the one teams forget most often, even though it is what tells you whether training is making progress. Like the other three, it must be tied to the job and the account.
Revised on 2 October 2026: Blackwell generation added (NVLink 5, InfiniBand XDR), IBM Storage Scale instead of Spectrum Scale, storage throughput aligned with the introductory course, interconnect latency restated as in the order of a microsecond.