Technical
HPC observability foundations
For: engineers and SREs · architectsPrerequisites: Linux command line, basics of Prometheus or Grafana.
Tuesday, 02:42. The training run just died.
Section titled “Tuesday, 02:42. The training run just died.”Illustrative scenario, built for the course.
Four hours of compute on eight A100 GPUs, gone: 32 GPU-hours, to multiply by your provider’s hourly price or your internal cost. Prices change fast and vary by contract: this guide gives none; use your provider’s dated price list.
02:42:14 INFO job 184729 step 14 done02:42:15 INFO rank 0 allreduce init02:42:16 ERROR ECC error GPU 302:42:17 WARN retrying step 1502:42:18 WARN retrying step 1502:42:19 INFO job 184729 terminated, exit 137Could you tell, in five minutes, which GPU failed, on which node, for which user? Whether the hardware was already degrading? Which account just absorbed those 32 GPU-hours?
Without a correlation prepared in advance, answering means cross-checking job logs, GPU metrics and Slurm accounting by hand. The data exists, but it sits in separate tools that nobody has wired together.
Three converging forces
Section titled “Three converging forces”| Force | Why observability becomes necessary |
|---|---|
| Scale of AI training | a training run uses tens to thousands of GPUs for days or weeks (order of magnitude); a silent GPU degradation loses part of that compute |
| Regulatory pressure | NIS2 for the entities in scope; Directive (EU) 2023/1791 on energy efficiency, whose Article 12 requires data centres with at least 500 kW of installed IT power to report their energy performance every year; usage reports on compute allocations. Auditors ask who did what, when |
| Multi-tenant accountability | showback, chargeback and fairshare all depend on knowing which account each joule belongs to |
What HPC changes for metrics, logs and traces
Section titled “What HPC changes for metrics, logs and traces”Metrics, logs and traces are still the foundation. HPC adds three constraints. Cardinality is higher, since you track every GPU and every rank. Retention gets longer, because researchers routinely compare against runs from years ago. And every observation has to be tied to a job, a user and an account, otherwise it is useless for chargeback or audit.
Content: three modules and a capstone
Section titled “Content: three modules and a capstone”| Module | Content | Lab |
|---|---|---|
| M1. Physical anatomy of the cluster | the five components and what they expose: node_exporter and NUMA placement, DCGM exporter and GPU UUID labels, InfiniBand or RoCE port counters, Lustre, IBM Storage Scale (formerly Spectrum Scale), BeeGFS, Slurm jobs, accounts and partitions | L1: a first Python Redfish exporter, scraped by VictoriaMetrics (or Prometheus), shown in Grafana, with a thermal alert triggered by fault injection |
| M2. Workload telemetry | tying GPU metrics to jobs: the Slurm prolog and epilog write, for each allocated GPU, the job ID into a directory read by dcgm-exporter (option --hpc-job-mapping-dir, variable DCGM_HPC_JOB_MAPPING_DIR); metrics then carry an hpc_job label. The job is then linked to its user and account with sacct or a Slurm exporter, for one business answer: power consumed per account | L2: dcgm-exporter with job mapping and a Slurm exporter side by side, a PromQL join on the job ID attributing GPU power to accounts, an energy-budget alert. L3: detect contention between compute instances of the same MIG GPU instance, which share that instance’s memory and bandwidth, contention invisible to standard utilization metrics |
| M3. Integration | a real deliverable: for a GPU cluster of several hundred nodes shared by 6 accounts, one dashboard that passes the five-second test in front of a director, plus three alerts that only page when intervention is truly needed | Capstone: design (six panels max), build in Grafana and vmalert, validate against three injected faults, defend in five minutes |
Every lab will include a starter file, a reference solution, automated validators and fault-injection scripts.
Objectives
Section titled “Objectives”- Deploy a standard stack from scratch: VictoriaMetrics (or Prometheus), Grafana, vmalert, Alertmanager.
- Read DCGM and Slurm metrics fluently, and know the five GPU and six Slurm metrics that matter most.
- Write a small Python Prometheus exporter that respects naming conventions and label discipline.
- Tie metrics from different collectors to a job, then to a Slurm account, to attribute consumption.
- Design a dashboard that passes the five-second test in front of a non-technical stakeholder.
- Write an alert rule that does not flap, with a runbook annotation and a documented threshold rationale.
Audience and prerequisites
Section titled “Audience and prerequisites”Intended for: SREs and platform engineers new to HPC, HPC sysadmins modernizing their stack, DevOps engineers onboarding AI workloads, research-computing staff under audit pressure, sovereign-cloud teams preparing AI infrastructure.
Out of scope: machine learning research, Slurm administration, Kubernetes-centric platforms.
Prerequisites: Linux command line, basic Prometheus or Grafana. Complete beginner? Start with Understanding HPC.
Planned follow-ups
Section titled “Planned follow-ups”| Follow-up | Content |
|---|---|
| Instrumentation, fabric, storage | write a production-grade Slurm exporter in Go, with a lab target of sustaining a simulated load of 5,000 jobs per second; InfiniBand fabric observability, Lustre in depth, capacity and queue-wait forecasting, full showback |
| AI training, sovereign stack, AIOps | rank-wise anomaly detection on a simulated 64-rank training run with an autoencoder, to catch an injected straggler whose penalty the lab sets (for example 30% of throughput, a simulation parameter, not a measurement); collective communication health, air-gapped stack, failure prediction |
Revised on 2 October 2026: prices removed, scenario cost expressed in GPU-hours, GPU-to-job attribution described with dcgm-exporter job mapping, MIG contention clarified (compute instances of the same GPU instance), reference to Directive (EU) 2023/1791, IBM Storage Scale, actual state of the labs.