Technical
LLM observability: the labs
For: engineers and SREs · architectsPrerequisites: Have already called an LLM API, know a classic observability tool and be able to read a Python script.
Scope: the LLM and RAG, practiced in labs; agents and MCP are covered in the article Observing an LLM system and in the GenAI observability method. Prerequisite: the concepts of the guide Understanding generative AI observability (signals, maturity, evaluators).
Classic tools watch infrastructure (CPU, memory, latency, HTTP error rates), but nothing they measure touches the meaning of what your model says. An HR assistant quoting an outdated leave policy, or a sales assistant inventing the price of a discontinued product, moves no dashboard at all. Infrastructure stays green while the model drifts.
The course is written for platform and data teams. It shows how to observe what the model produces, spot the three families of drift, keep cost under control and set a quality level in production that can actually be measured.
What you will be able to do after the labs
Section titled “What you will be able to do after the labs”- tell LLM observability apart from classic application observability;
- instrument an LLM application with OpenTelemetry GenAI semantic conventions;
- build quality, cost and inference performance dashboards;
- detect data, concept and pipeline drift with reproducible indicators;
- set up continuous evaluation (LLM-as-judge, RAG metrics, user feedback);
- define SLOs and alerts suited to generative workloads;
- deploy it all on a self-hosted stack where every component can be replaced.
Program
Section titled “Program”| Step | Content | Indicative duration |
|---|---|---|
| Module 1 | why an LLM cannot be observed like a classic application | 1 h 30 |
| Module 2 | OpenTelemetry GenAI conventions | 1 h 45 |
| Lab 1 | instrument an LLM application | 1 h 45 |
| Lab 2 | an end-to-end observed RAG pipeline | 1 h 30 |
| Module 3 | the three drifts | 1 h 30 |
| Module 4 | continuous evaluation in production | 1 h 45 |
| Lab 3 | detect drift automatically | 1 h 30 |
| Lab 4 | cost, performance, SLOs and alerts | 1 h 35 |
| Module 5 | industrialize: 30-, 60- and 90-day plan, AI Act | 30 min |
Prerequisites
Section titled “Prerequisites”- having already called an LLM API (Mistral, OpenAI, Claude) or a local model through Ollama;
- knowing at least one classic observability tool (Prometheus, Grafana, Elastic);
- being able to read and modify a Python script;
- ideally, knowing the basics of OpenTelemetry.
The stack
Section titled “The stack”| Component | Kit version | License | Role |
|---|---|---|---|
| OpenTelemetry (Python SDK and Collector contrib) | SDK 1.28.2, Collector 0.108.0 | Apache 2.0 | instrumentation, GenAI conventions, routing |
| VictoriaMetrics | v1.110.0 | Apache 2.0 | metrics storage |
| Tempo | 2.6.0 | AGPL 3.0 | trace storage |
| Grafana | 11.2.0 | AGPL 3.0 | dashboards, exploration, alerts |
| Arize Phoenix | 5.9.1 | Elastic License 2.0 (source-available) | conversation exploration, evaluation |
| Qdrant | v1.12.0 | Apache 2.0 | RAG vector database |
| Ollama | 0.3.12 | MIT | local LLM (Mistral 7B) and embedding model |
Every component is open source except Phoenix, whose Elastic License 2.0 makes the code available without being an open source license. The whole stack runs on a workstation with Docker (16 GB of memory recommended). Each component has equivalents, open source or commercial: module 1 gives an overview and chapter 6 of the guide maps the full tools landscape and their licenses.
Lab kit status
Section titled “Lab kit status”The kit (folder labs/llm-observability/, Apache 2.0 license) is not published yet: it will be on the Labs & Trainings forge. Until then, the lab lessons read as exercise statements.
On 2 October 2026 the kit was fixed on the following points, which prevented it from producing what the lessons describe:
- unpinned Python dependencies, while the recent Qdrant client no longer has the
search()method used by labs 2 to 4: dependencies pinned and code moved toquery_points(); - metric names missing the
_secondsand_totalsuffixes added by the exporter: empty latency and cost panels, queries fixed; - Grafana data sources without a
uid: the five alerts and the trace-to-metric link pointed to a non-existent source; - a Collector cost calculation that never applied and a filter defined outside any pipeline: removed, cost is computed in the application;
- Phoenix at the center of the diagram but with no export: a trace pipeline to Phoenix was added;
- in-house extensions under
gen_ai.*and the deprecatedgen_ai.systemattribute: moved tomttl.*andgen_ai.provider.name; - promised metrics that were missing (hallucinations, tenant on the drift score, per-feature breakdown): added;
- an impossible concept drift scenario (reindexing at every run) and alerts that lab 4 could not trigger: explicit reindexing and a configurable simulator;
- an illustrative price grid in lab 4: removed. Prices change fast and vary by contract: this guide gives none; use your provider’s dated price list. The lab budget and SLOs are in tokens, and the euro cost is computed only with a dated price list you supply.
Revised on 2 October 2026: total duration corrected (13 hours), prerequisites, stack versions and licenses added, kit status and list of fixes added, AI Act described in terms of monitoring and logs; prices removed.
Revised on 4 October 2026: title “LLM observability: the labs”, scope sentence (LLM and RAG in labs; agents and MCP referred to the article and the method; prerequisite: the guide), progression box added, page moved to MDX.
Revised on 5 October 2026: lab 4 duration aligned with its steps (1 h 35), bringing the total to about 13 h 20.