Technical
Lab 3: detect drift automatically
For: engineers and SREsPrerequisites: Have completed lab 2 and read module 3 of the course.
Duration: 1 h 30. Prerequisites: lab 2 completed, module 3.
The scenario
Section titled “The scenario”The script runs four phases in sequence and leaves, in the score, the mark of a deliberately injected drift. It publishes under its own service, lab3-drift, reusing the lab 2 pipeline. Detection: the phase average falls below 0.6, the threshold shared by all labs.
| Phase | What happens | Tracked score | Expected order of magnitude |
|---|---|---|---|
| Baseline | 5 in-domain questions, 3 repetitions | faithfulness to context | above the threshold, around 0.75 |
| Data drift | 5 out-of-domain questions (cooking, weather, sport), 2 repetitions | faithfulness to context | well below the threshold, around 0.20 |
| Concept drift | 2 questions whose answer has changed in reality, 3 repetitions | comparison with ground truth | close to 0 as long as the knowledge base is not updated |
| Pipeline drift | top_k silently drops from 3 to 1 | faithfulness to context | below the threshold, around 0.45 |
These values are indicative: they vary with the model and from one run to the next. The kit’s faithfulness is a ratio of words shared between the answer and the retrieved context, not a semantic measure; it is enough to show a break, not to grade an answer finely.
Step 1: run the four phases (10 min)
Section titled “Step 1: run the four phases (10 min)”cd app && python lab3_drift_detection.py# a single phase, several passes:python lab3_drift_detection.py --phase data --repeat 2Step 2: visualize the break (20 min)
Section titled “Step 2: visualize the break (20 min)”On the Drift Detection dashboard, look for:
- the score drop when moving from the baseline to data drift;
- the drift events emitted (
mttl_drift_events_total, one per drifting phase); - the table of drifting customers (
drift-tenant, average score below 0.6 over 15 min).
For pipeline drift, open a trace: the cause, mttl.rag.top_k = 1, is visible in the attributes of the retrieval span.
Step 3: concept drift (30 min)
Section titled “Step 3: concept drift (30 min)”The concept phase compares the answers with a ground truth (VERITE_TERRAIN in the script): the fictional offer now costs 59 EUR and responds within 2 hours, while the knowledge base still says 49 EUR and 4 hours. Faithfulness to context stays high, since the model faithfully repeats an outdated knowledge base; only the comparison with ground truth reveals the drift.
- Edit
DOCSinlab2_rag_pipeline.py(59 EUR, 2 hours). - Run
python lab3_drift_detection.py --phase conceptagain: nothing changes, because the scripts never reindex on their own. - Run
python lab2_rag_pipeline.py --reindex, then theconceptphase again: the score goes back up.
Question for discussion: how do you industrialize this detection? Versioning the knowledge base, rebuilding embeddings, non-regression tests on a reference set.
Step 4: Grafana or Phoenix (20 min)
Section titled “Step 4: Grafana or Phoenix (20 min)”Traces also go to Phoenix: the Collector sends them through the otlp/phoenix exporter (port 4317 inside the Docker network, published on 4319 on the host). Open http://localhost:6006 and compare the two approaches.
# otel/otel-collector-config.yaml (excerpt)exporters: otlp/phoenix: endpoint: phoenix:4317 tls: insecure: trueservice: pipelines: traces: receivers: [otlp] processors: [resource/defaults, transform/genai_compat, batch] exporters: [otlp/tempo, otlp/phoenix]| Grafana | Phoenix | |
|---|---|---|
| Strength | metrics, alerts, on-call | conversation exploration, evaluation |
| Use | detect and alert | understand and qualify |
So you keep both. The alert comes from Grafana, then the investigation of the conversations involved happens in Phoenix.
Step 5: response plan (10 min)
Section titled “Step 5: response plan (10 min)”Write a plan for each type of drift:
- data: whom to notify, and how quickly;
- concept: knowledge base update process, frequency;
- pipeline: non-regression tests and quality gates in continuous integration.
Lab recap
Section titled “Lab recap”| Notion | What the lab shows |
|---|---|
| Statistical detection | moving average and threshold: enough to get started |
| Next step | KS, PSI, Wasserstein |
| Grafana and Phoenix | operational alerting on one side, exploration and evaluation on the other |
| Triage | every alert has an owner and a procedure |
Revised on 2 October 2026: box on the kit status, expected scores presented as indicative, the kit’s faithfulness measure explained, customer label added to the drift metric, reindexing separated into the --reindex option, configuration of trace export to Phoenix.