Technical
Beyond the LLM: GPUs and model quality
For: engineers and SREs · architectsPrerequisites: Having completed 'LLM observability: the labs'; experience industrializing AI models in production.
This path is the last step of the progression: it goes beyond the LLM, towards GPU infrastructure and then model quality over time. It targets ML engineers, platform engineers, SREs and MLOps teams already industrializing models.
Prerequisite: the LLM observability: the labs course (GenAI instrumentation, RAG, drift, cost, alerting), which this path assumes and does not repeat. For the concepts, see the guide.
Its own scope: classic ML, statistical drift (PSI, KS, Wasserstein), the model registry (MLflow) and retraining belong to this path; the guide, focused on generative AI, leaves them out of its scope.
- GPU and AI infrastructure monitoring: DCGM, NVIDIA exporters, Slurm, Kubernetes. Module in preparation; HPC observability foundations and the HPC AI reference architecture already cover part of it.
- Advanced drift detection and model quality: PSI, KS, Wasserstein, Evidently, MLflow, retraining. Module in preparation.
Module 1: GPU and AI infrastructure monitoring
Section titled “Module 1: GPU and AI infrastructure monitoring”On GPU infrastructure, the utilization figure tells you very little. A GPU showing 80% utilization can hide saturated memory bandwidth, underused tensor cores or a network bottleneck between nodes. Illustrative scenario: a training run shows 85% utilization while memory bandwidth is at 99%; memory is the bottleneck, and the utilization figure does not show it.
Objectives
- map the layers of an AI infrastructure: GPU, memory, interconnect, scheduler, application;
- deploy the DCGM exporter and collect advanced NVIDIA metrics: SM occupancy, tensor core utilization, memory bandwidth, thermal throttling, NVLink and NCCL, power;
- instrument Slurm jobs and Kubernetes AI pods with the right labels;
- build a dashboard usable by an SRE who is not an AI specialist;
- detect problem patterns: saturated bandwidth, slow NCCL, insufficient memory during training;
- define the SLOs of an AI infrastructure: availability, memory overrun rate, wait time, throughput.
Planned labs (in preparation): a DCGM stack with a Prometheus-compatible metrics backend, detecting underused tensor cores on a synthetic job; then a monitored Slurm or Kubernetes cluster with job-level metrics, failure rate and cost allocation per team.
Module 2: advanced drift detection and model quality
Section titled “Module 2: advanced drift detection and model quality”Spotting drift is not enough. You still have to qualify it statistically, tell a real change from a false positive, choose between alerting, retraining or letting it pass, and keep a record of those decisions over time.
Objectives
- formally distinguish drift types and place them in the site’s three families (data, concept, pipeline), whose reference is module 3 of the LLM course: covariate shift, prior shift and virtual drift belong to data, concept drift to concept; then their patterns: sudden, gradual, incremental, recurring;
- implement PSI, Kolmogorov-Smirnov, Chi², Wasserstein and MMD on tabular features and embeddings, knowing the strengths and limits of each;
- configure Evidently or an equivalent in production;
- define a retraining strategy triggered by drift, performance or calendar, with champion and challenger and safe rollback;
- keep a model registry with MLflow or an equivalent, and production lineage;
- set up an AI quality committee and its rituals, and prepare the technical documentation the AI Act requires for high-risk systems. Since Regulation (EU) 2026/1744, these obligations apply from 2 December 2027 for Annex III systems and from 2 August 2028 for Annex I systems.
Planned labs (in preparation): detection on tabular features with seasonal false alarms; detection on text embeddings inside a RAG pipeline; a full chain from Evidently alert to webhook, MLflow job, validation and controlled deployment.
Revised on 2 October 2026: AI Act timeline updated after Regulation (EU) 2026/1744, modules 2 and 3 and their labs flagged as in preparation, GPU example presented as an illustrative scenario, mention of a non-existent security path removed.
Revised on 4 October 2026: path renamed “Beyond the LLM: GPUs and model quality”, the LLM course becomes a prerequisite rather than a module, modules renumbered, drift terms placed in the site’s taxonomy (data, concept, pipeline) with a link to module 3 of the LLM course, own scope clarified (classic ML, MLflow, retraining), page moved to MDX with the progression box.