Skip to content

TechnicalPractitioner

Module 3: the three drifts

For: engineers and SREsPrerequisites: Have followed lab 2 of the course on the RAG pipeline.

A model in production degrades in three ways, all silent and invisible from infrastructure monitoring.

DriftCauseExample
Datausers ask different questionsan HR assistant designed for leave receives questions about internal mobility
Conceptexternal reality has changed without the system knowinga collective agreement changes, the assistant still quotes the old version
Pipelinea component was changed without testingtop_k goes from 5 to 1 during a code cleanup, the context becomes insufficient

Pipeline drift does not only come from your own team. When you call an LLM through an API, the provider can change the model served behind the same name or alias without your code moving: that is pipeline drift too. To see it, record in each span the requested model (gen_ai.request.model) and the model the provider actually returned (gen_ai.response.model), along with the prompt template version and the RAG chunking settings.

From the cause to the signal that reveals it, then to the matching response:

flowchart LR
  C1["Data<br/>questions change"] --> S1["Intent distribution,<br/>out-of-domain volume"] --> A1["Extend the document base<br/>and the reference set"]
  C2["Concept<br/>reality changes"] --> S2["Faithfulness against ground truth,<br/>negative feedback"] --> A2["Update the sources<br/>and reindex"]
  C3["Pipeline<br/>a component changes"] --> S3["Span attributes,<br/>latency per step"] --> A3["Roll back the change<br/>and test it"]

In my view, pipeline drift is the easiest to diagnose, provided the pipeline configuration has been instrumented in the spans: its cause is found in the trace itself, without having to analyze user questions.

This module is the site’s reference for the drift taxonomy: the guide’s glossary and the other pieces align with it. The terms you will meet elsewhere fit into the three families as follows:

Term encounteredCanonical family
input drift, covariate shift, embedding drift (detection)data
prior shift (the label distribution changes)data
virtual drift (the input distribution changes without changing the input-to-output relationship)data
conceptual drift, concept driftconcept
silent change of the provider or the model, of the prompt template, of RAG chunkingpipeline
output driftnot a family: it is the symptom revealed by evaluations, whose cause is data, concept or pipeline drift
DriftIndicatorMethodTrigger
Dataintent distributionPSI or Kolmogorov-Smirnov on embeddingsPSI or KS statistic above a threshold calibrated on a reference period
Dataout-of-domain volumeintent classifier or similarity with the basemore than 15% of traffic out of domain
Conceptfaithfulness against ground truthLLM judge on a sampledrop of more than 10 points
Conceptnegative feedbackexplicit thumbs-down or abandonmentmore than 5% over 24 h
Pipelinedistribution of span attributeshistogram of top_k, of the model nameappearance of a new value
Pipelinelatency per step7-day sliding histogramunexplained variation of more than 30%

These thresholds are illustrative starting points, not drawn from a study: recalibrate them on your own traffic after a few weeks. The 0.6 quality score is the one used by every lab in the kit (lab 3 detection, lab 4 alert and SLO).

The distinctive sign of a drift is decorrelation: all technical indicators are normal, only the quality score drops. A dashboard that puts latency, error rate and quality score side by side makes this decorrelation visible at a glance.

A moving average of the quality score and a threshold are enough to start. Statistical tests (PSI, KS, Wasserstein) come later, when volume justifies the effort. Advanced statistical detection is covered in the path Beyond the LLM: GPUs and model quality.

Data, concept and pipeline drift have neither the same cause nor the same response. Pipeline drift can be read in span attributes, provided the configuration was put there. In every case, the symptom to watch is the same: green infrastructure and a dropping quality score, measured against the retrieved context or, for concept drift, against ground truth (faithfulness to the context, for its part, can stay high). To start, a moving average and a threshold are enough.

Revised on 2 October 2026: unsourced claim about the frequency of pipeline drift removed, thresholds presented as illustrative and aligned on 0.6, module code replaced by a link to the advanced AI observability path.

Revised on 4 October 2026: module designated as the site’s reference for the drift taxonomy, mapping table of terms added (output drift presented as a symptom), silent change of the provider or model classified as pipeline drift.