Technical
Module 3: the three drifts
For: engineers and SREsPrerequisites: Have followed lab 2 of the course on the RAG pipeline.
A model in production degrades in three ways, all silent and invisible from infrastructure monitoring.
Data, concept, pipeline
Section titled “Data, concept, pipeline”| Drift | Cause | Example |
|---|---|---|
| Data | users ask different questions | an HR assistant designed for leave receives questions about internal mobility |
| Concept | external reality has changed without the system knowing | a collective agreement changes, the assistant still quotes the old version |
| Pipeline | a component was changed without testing | top_k goes from 5 to 1 during a code cleanup, the context becomes insufficient |
Pipeline drift does not only come from your own team. When you call an LLM through an API, the provider can change the model served behind the same name or alias without your code moving: that is pipeline drift too. To see it, record in each span the requested model (gen_ai.request.model) and the model the provider actually returned (gen_ai.response.model), along with the prompt template version and the RAG chunking settings.
From the cause to the signal that reveals it, then to the matching response:
flowchart LR C1["Data<br/>questions change"] --> S1["Intent distribution,<br/>out-of-domain volume"] --> A1["Extend the document base<br/>and the reference set"] C2["Concept<br/>reality changes"] --> S2["Faithfulness against ground truth,<br/>negative feedback"] --> A2["Update the sources<br/>and reindex"] C3["Pipeline<br/>a component changes"] --> S3["Span attributes,<br/>latency per step"] --> A3["Roll back the change<br/>and test it"]
In my view, pipeline drift is the easiest to diagnose, provided the pipeline configuration has been instrumented in the spans: its cause is found in the trace itself, without having to analyze user questions.
Mapping of terms
Section titled “Mapping of terms”This module is the site’s reference for the drift taxonomy: the guide’s glossary and the other pieces align with it. The terms you will meet elsewhere fit into the three families as follows:
| Term encountered | Canonical family |
|---|---|
| input drift, covariate shift, embedding drift (detection) | data |
| prior shift (the label distribution changes) | data |
| virtual drift (the input distribution changes without changing the input-to-output relationship) | data |
| conceptual drift, concept drift | concept |
| silent change of the provider or the model, of the prompt template, of RAG chunking | pipeline |
| output drift | not a family: it is the symptom revealed by evaluations, whose cause is data, concept or pipeline drift |
Reproducible indicators
Section titled “Reproducible indicators”| Drift | Indicator | Method | Trigger |
|---|---|---|---|
| Data | intent distribution | PSI or Kolmogorov-Smirnov on embeddings | PSI or KS statistic above a threshold calibrated on a reference period |
| Data | out-of-domain volume | intent classifier or similarity with the base | more than 15% of traffic out of domain |
| Concept | faithfulness against ground truth | LLM judge on a sample | drop of more than 10 points |
| Concept | negative feedback | explicit thumbs-down or abandonment | more than 5% over 24 h |
| Pipeline | distribution of span attributes | histogram of top_k, of the model name | appearance of a new value |
| Pipeline | latency per step | 7-day sliding histogram | unexplained variation of more than 30% |
These thresholds are illustrative starting points, not drawn from a study: recalibrate them on your own traffic after a few weeks. The 0.6 quality score is the one used by every lab in the kit (lab 3 detection, lab 4 alert and SLO).
Green infrastructure, degraded quality
Section titled “Green infrastructure, degraded quality”The distinctive sign of a drift is decorrelation: all technical indicators are normal, only the quality score drops. A dashboard that puts latency, error rate and quality score side by side makes this decorrelation visible at a glance.
Start simple
Section titled “Start simple”A moving average of the quality score and a threshold are enough to start. Statistical tests (PSI, KS, Wasserstein) come later, when volume justifies the effort. Advanced statistical detection is covered in the path Beyond the LLM: GPUs and model quality.
In summary
Section titled “In summary”Data, concept and pipeline drift have neither the same cause nor the same response. Pipeline drift can be read in span attributes, provided the configuration was put there. In every case, the symptom to watch is the same: green infrastructure and a dropping quality score, measured against the retrieved context or, for concept drift, against ground truth (faithfulness to the context, for its part, can stay high). To start, a moving average and a threshold are enough.
Revised on 2 October 2026: unsourced claim about the frequency of pipeline drift removed, thresholds presented as illustrative and aligned on 0.6, module code replaced by a link to the advanced AI observability path.
Revised on 4 October 2026: module designated as the site’s reference for the drift taxonomy, mapping table of terms added (output drift presented as a symptom), silent change of the provider or model classified as pipeline drift.