Skip to content

TechnicalBeginner

HPC and AI

For: engineers and SREs · architects · team managers · executives and CIOsPrerequisites: Have read lesson 2 of the course.

Before (until around 2018)Today
HPC: processor clusters, double-precision simulation, MPIthousands of GPUs and a very fast interconnect
AI: a few GPUs, small datasetsshared parallel storage and scheduler
separate hardware, tools and conferencescomparable observability and operations

Large-scale AI therefore runs on the same foundation as classic HPC. Mostly the terms and the tracked indicators change.

Numerical precision. Classic HPC often requires double precision (FP64) for stability. AI tolerates 16 bits, even 8 bits: GPUs go faster, but precision and reproducibility issues appear.

Data scale. Training sets reach hundreds of terabytes. Feeding thousands of GPUs without interruption is a storage and network challenge in its own right.

Observability. AI adds its own indicators: loss curve, gradient norms, step time, per-rank GPU utilization, energy per processed token. Observing a training run is becoming a specialty.

An order of magnitude, as an illustration: a 7-day training run on 64 GPUs consumes 64 x 7 x 24 = 10,752 GPU-hours. Multiply this volume by your provider’s hourly price or your internal cost. The same calculation applied to the largest training runs, on thousands of GPUs for weeks, gives hundreds of thousands, even millions of GPU-hours. A silent degradation of a single GPU that loses a checkpoint can cost hours, even days of computation, depending on checkpoint frequency. At this level of consumption, observability becomes an economic question as much as a technical one.

Prices change fast and vary by contract: this guide gives none; use your provider’s dated price list.

You now know the five components of a cluster (login node, scheduler, compute nodes, interconnect, parallel storage) and the three patterns of parallelism (trivial, tight coupling, data parallelism). You also know why large-scale AI rests on the same foundation as classic HPC.

  • Read the first pages of the Slurm sbatch documentation: the abstractions become concrete.
  • Browse the Top500 ranking of supercomputers.
  • Try an MPI or OpenMP program on your workstation: you will not break any record, but you will get a feel for parallelism.
  • Continue with HPC observability foundations.

Revised on 2 October 2026: cost of a training run recalculated explicitly and presented as illustrative (64 GPUs for 7 days, or 10,752 GPU-hours). This calculation replaces the mention “hundreds of thousands of euros”; prices removed.