Technical
HPC and AI
For: engineers and SREs · architects · team managers · executives and CIOsPrerequisites: Have read lesson 2 of the course.
Two worlds that became one
Section titled “Two worlds that became one”| Before (until around 2018) | Today |
|---|---|
| HPC: processor clusters, double-precision simulation, MPI | thousands of GPUs and a very fast interconnect |
| AI: a few GPUs, small datasets | shared parallel storage and scheduler |
| separate hardware, tools and conferences | comparable observability and operations |
Large-scale AI therefore runs on the same foundation as classic HPC. Mostly the terms and the tracked indicators change.
What still sets AI apart from classic HPC
Section titled “What still sets AI apart from classic HPC”Numerical precision. Classic HPC often requires double precision (FP64) for stability. AI tolerates 16 bits, even 8 bits: GPUs go faster, but precision and reproducibility issues appear.
Data scale. Training sets reach hundreds of terabytes. Feeding thousands of GPUs without interruption is a storage and network challenge in its own right.
Observability. AI adds its own indicators: loss curve, gradient norms, step time, per-rank GPU utilization, energy per processed token. Observing a training run is becoming a specialty.
What it changes for operations
Section titled “What it changes for operations”An order of magnitude, as an illustration: a 7-day training run on 64 GPUs consumes 64 x 7 x 24 = 10,752 GPU-hours. Multiply this volume by your provider’s hourly price or your internal cost. The same calculation applied to the largest training runs, on thousands of GPUs for weeks, gives hundreds of thousands, even millions of GPU-hours. A silent degradation of a single GPU that loses a checkpoint can cost hours, even days of computation, depending on checkpoint frequency. At this level of consumption, observability becomes an economic question as much as a technical one.
Prices change fast and vary by contract: this guide gives none; use your provider’s dated price list.
Recap of the three lessons
Section titled “Recap of the three lessons”You now know the five components of a cluster (login node, scheduler, compute nodes, interconnect, parallel storage) and the three patterns of parallelism (trivial, tight coupling, data parallelism). You also know why large-scale AI rests on the same foundation as classic HPC.
Going further
Section titled “Going further”- Read the first pages of the Slurm
sbatchdocumentation: the abstractions become concrete. - Browse the Top500 ranking of supercomputers.
- Try an MPI or OpenMP program on your workstation: you will not break any record, but you will get a feel for parallelism.
- Continue with HPC observability foundations.
Revised on 2 October 2026: cost of a training run recalculated explicitly and presented as illustrative (64 GPUs for 7 days, or 10,752 GPU-hours). This calculation replaces the mention “hundreds of thousands of euros”; prices removed.