Skip to content

TechnicalPractitioner

HPC AI cluster architecture

For: engineers and SREs · architectsPrerequisites: Basic HPC notions: node, scheduler, interconnect.

flowchart TB
  SCHED["Slurm / PBS scheduler"]
  subgraph N1["GPU node 1"]
    direction LR
    G0["GPU 0<br/>80 GB HBM"] --- G1["GPU 1"] --- G2["GPU 2"] --- G3["GPU 3"]
  end
  subgraph N2["GPU node 2"]
    direction LR
    G4["GPU 4"] --- G5["GPU 5"] --- G6["GPU 6"] --- G7["GPU 7"]
  end
  IB["InfiniBand / RoCE<br/>latency in the order of a microsecond"]
  FS["Parallel storage<br/>Lustre / IBM Storage Scale / BeeGFS"]
  SCHED --> N1
  SCHED --> N2
  N1 <--> IB
  N2 <--> IB
  IB <--> FS

Inside a node, GPUs talk over NVLink; between nodes, over InfiniBand or RoCE. Parallel storage feeds everyone, and the scheduler decides what runs where.

LayerWhat it containsWhat to observeWhat to protect
L1. GPU nodes4 to 8 GPUs per node linked by NVLink (600 GB/s on A100, 900 GB/s on H100, 1.8 TB/s on Blackwell), 80 GB HBM per GPU on A100 and H100, more on later generations (141 GB on H200), 2 orchestration CPUs, hundreds of GB of RAM, local NVMe for stagingdcgm-exporter: SM occupancy, HBM bandwidth, ECC errors, throttling, powerfirmware, GPU memory between jobs
L2. InterconnectInfiniBand HDR (200 Gb/s), NDR (400 Gb/s) or XDR (800 Gb/s, Blackwell generation), or RoCE; non-blocking fat-tree topology; gradients flow here during NCCL AllReduceibstat, perfquery, port counters, retransmitsthe RDMA blind spot
L3. Parallel storageLustre or IBM Storage Scale (formerly Spectrum Scale, or GPFS), files spread over dozens of servers, aggregate throughput from tens of GB/s to several TB/s depending on system size; GPUs must never wait for dataLustre throughput, metadata latencytraining data integrity
L4. SchedulerSlurm or PBS: queues, GPU allocation, job isolation; prolog and epilog scripts to inject instrumentationslurm-exporter: jobs, queues, wait timeisolation between jobs and tenants
L5. ML software stackCUDA or ROCm driver, PyTorch or JAX, NCCL for collective communication, Docker or Apptainer (formerly Singularity) containersPyTorch profiler, NCCL logssoftware supply chain
FamilyExamples
GPUSM utilization, HBM bandwidth, ECC errors, clock throttling
NetworkInfiniBand port counters, retransmit rate, per-link throughput
StorageLustre throughput, metadata latency
Modelloss, gradient norm, step time, AllReduce latency

The fourth family is the one teams forget most often, even though it is what tells you whether training is making progress. Like the other three, it must be tied to the job and the account.

Revised on 2 October 2026: Blackwell generation added (NVLink 5, InfiniBand XDR), IBM Storage Scale instead of Spectrum Scale, storage throughput aligned with the introductory course, interconnect latency restated as in the order of a microsecond.