Skip to content

TechnicalExpert

Beyond the LLM: GPUs and model quality

For: engineers and SREs · architectsPrerequisites: Having completed 'LLM observability: the labs'; experience industrializing AI models in production.

Reading mode

This path is the last step of the progression: it goes beyond the LLM, towards GPU infrastructure and then model quality over time. It targets ML engineers, platform engineers, SREs and MLOps teams already industrializing models.

Prerequisite: the LLM observability: the labs course (GenAI instrumentation, RAG, drift, cost, alerting), which this path assumes and does not repeat. For the concepts, see the guide.

Its own scope: classic ML, statistical drift (PSI, KS, Wasserstein), the model registry (MLflow) and retraining belong to this path; the guide, focused on generative AI, leaves them out of its scope.

  1. GPU and AI infrastructure monitoring: DCGM, NVIDIA exporters, Slurm, Kubernetes. Module in preparation; HPC observability foundations and the HPC AI reference architecture already cover part of it.
  2. Advanced drift detection and model quality: PSI, KS, Wasserstein, Evidently, MLflow, retraining. Module in preparation.

Module 1: GPU and AI infrastructure monitoring

Section titled “Module 1: GPU and AI infrastructure monitoring”

On GPU infrastructure, the utilization figure tells you very little. A GPU showing 80% utilization can hide saturated memory bandwidth, underused tensor cores or a network bottleneck between nodes. Illustrative scenario: a training run shows 85% utilization while memory bandwidth is at 99%; memory is the bottleneck, and the utilization figure does not show it.

Objectives

  • map the layers of an AI infrastructure: GPU, memory, interconnect, scheduler, application;
  • deploy the DCGM exporter and collect advanced NVIDIA metrics: SM occupancy, tensor core utilization, memory bandwidth, thermal throttling, NVLink and NCCL, power;
  • instrument Slurm jobs and Kubernetes AI pods with the right labels;
  • build a dashboard usable by an SRE who is not an AI specialist;
  • detect problem patterns: saturated bandwidth, slow NCCL, insufficient memory during training;
  • define the SLOs of an AI infrastructure: availability, memory overrun rate, wait time, throughput.

Planned labs (in preparation): a DCGM stack with a Prometheus-compatible metrics backend, detecting underused tensor cores on a synthetic job; then a monitored Slurm or Kubernetes cluster with job-level metrics, failure rate and cost allocation per team.

Module 2: advanced drift detection and model quality

Section titled “Module 2: advanced drift detection and model quality”

Spotting drift is not enough. You still have to qualify it statistically, tell a real change from a false positive, choose between alerting, retraining or letting it pass, and keep a record of those decisions over time.

Objectives

  • formally distinguish drift types and place them in the site’s three families (data, concept, pipeline), whose reference is module 3 of the LLM course: covariate shift, prior shift and virtual drift belong to data, concept drift to concept; then their patterns: sudden, gradual, incremental, recurring;
  • implement PSI, Kolmogorov-Smirnov, Chi², Wasserstein and MMD on tabular features and embeddings, knowing the strengths and limits of each;
  • configure Evidently or an equivalent in production;
  • define a retraining strategy triggered by drift, performance or calendar, with champion and challenger and safe rollback;
  • keep a model registry with MLflow or an equivalent, and production lineage;
  • set up an AI quality committee and its rituals, and prepare the technical documentation the AI Act requires for high-risk systems. Since Regulation (EU) 2026/1744, these obligations apply from 2 December 2027 for Annex III systems and from 2 August 2028 for Annex I systems.

Planned labs (in preparation): detection on tabular features with seasonal false alarms; detection on text embeddings inside a RAG pipeline; a full chain from Evidently alert to webhook, MLflow job, validation and controlled deployment.

Revised on 2 October 2026: AI Act timeline updated after Regulation (EU) 2026/1744, modules 2 and 3 and their labs flagged as in preparation, GPU example presented as an illustrative scenario, mention of a non-existent security path removed.

Revised on 4 October 2026: path renamed “Beyond the LLM: GPUs and model quality”, the LLM course becomes a prerequisite rather than a module, modules renumbered, drift terms placed in the site’s taxonomy (data, concept, pipeline) with a link to module 3 of the LLM course, own scope clarified (classic ML, MLflow, retraining), page moved to MDX with the progression box.