Skip to content

TechnicalPractitioner

LLM observability: the labs

For: engineers and SREs · architectsPrerequisites: Have already called an LLM API, know a classic observability tool and be able to read a Python script.

Reading mode

Scope: the LLM and RAG, practiced in labs; agents and MCP are covered in the article Observing an LLM system and in the GenAI observability method. Prerequisite: the concepts of the guide Understanding generative AI observability (signals, maturity, evaluators).

Classic tools watch infrastructure (CPU, memory, latency, HTTP error rates), but nothing they measure touches the meaning of what your model says. An HR assistant quoting an outdated leave policy, or a sales assistant inventing the price of a discontinued product, moves no dashboard at all. Infrastructure stays green while the model drifts.

The course is written for platform and data teams. It shows how to observe what the model produces, spot the three families of drift, keep cost under control and set a quality level in production that can actually be measured.

What you will be able to do after the labs

Section titled “What you will be able to do after the labs”
  • tell LLM observability apart from classic application observability;
  • instrument an LLM application with OpenTelemetry GenAI semantic conventions;
  • build quality, cost and inference performance dashboards;
  • detect data, concept and pipeline drift with reproducible indicators;
  • set up continuous evaluation (LLM-as-judge, RAG metrics, user feedback);
  • define SLOs and alerts suited to generative workloads;
  • deploy it all on a self-hosted stack where every component can be replaced.
StepContentIndicative duration
Module 1why an LLM cannot be observed like a classic application1 h 30
Module 2OpenTelemetry GenAI conventions1 h 45
Lab 1instrument an LLM application1 h 45
Lab 2an end-to-end observed RAG pipeline1 h 30
Module 3the three drifts1 h 30
Module 4continuous evaluation in production1 h 45
Lab 3detect drift automatically1 h 30
Lab 4cost, performance, SLOs and alerts1 h 35
Module 5industrialize: 30-, 60- and 90-day plan, AI Act30 min
  • having already called an LLM API (Mistral, OpenAI, Claude) or a local model through Ollama;
  • knowing at least one classic observability tool (Prometheus, Grafana, Elastic);
  • being able to read and modify a Python script;
  • ideally, knowing the basics of OpenTelemetry.
ComponentKit versionLicenseRole
OpenTelemetry (Python SDK and Collector contrib)SDK 1.28.2, Collector 0.108.0Apache 2.0instrumentation, GenAI conventions, routing
VictoriaMetricsv1.110.0Apache 2.0metrics storage
Tempo2.6.0AGPL 3.0trace storage
Grafana11.2.0AGPL 3.0dashboards, exploration, alerts
Arize Phoenix5.9.1Elastic License 2.0 (source-available)conversation exploration, evaluation
Qdrantv1.12.0Apache 2.0RAG vector database
Ollama0.3.12MITlocal LLM (Mistral 7B) and embedding model

Every component is open source except Phoenix, whose Elastic License 2.0 makes the code available without being an open source license. The whole stack runs on a workstation with Docker (16 GB of memory recommended). Each component has equivalents, open source or commercial: module 1 gives an overview and chapter 6 of the guide maps the full tools landscape and their licenses.

The kit (folder labs/llm-observability/, Apache 2.0 license) is not published yet: it will be on the Labs & Trainings forge. Until then, the lab lessons read as exercise statements.

On 2 October 2026 the kit was fixed on the following points, which prevented it from producing what the lessons describe:

  • unpinned Python dependencies, while the recent Qdrant client no longer has the search() method used by labs 2 to 4: dependencies pinned and code moved to query_points();
  • metric names missing the _seconds and _total suffixes added by the exporter: empty latency and cost panels, queries fixed;
  • Grafana data sources without a uid: the five alerts and the trace-to-metric link pointed to a non-existent source;
  • a Collector cost calculation that never applied and a filter defined outside any pipeline: removed, cost is computed in the application;
  • Phoenix at the center of the diagram but with no export: a trace pipeline to Phoenix was added;
  • in-house extensions under gen_ai.* and the deprecated gen_ai.system attribute: moved to mttl.* and gen_ai.provider.name;
  • promised metrics that were missing (hallucinations, tenant on the drift score, per-feature breakdown): added;
  • an impossible concept drift scenario (reindexing at every run) and alerts that lab 4 could not trigger: explicit reindexing and a configurable simulator;
  • an illustrative price grid in lab 4: removed. Prices change fast and vary by contract: this guide gives none; use your provider’s dated price list. The lab budget and SLOs are in tokens, and the euro cost is computed only with a dated price list you supply.

Revised on 2 October 2026: total duration corrected (13 hours), prerequisites, stack versions and licenses added, kit status and list of fixes added, AI Act described in terms of monitoring and logs; prices removed.

Revised on 4 October 2026: title “LLM observability: the labs”, scope sentence (LLM and RAG in labs; agents and MCP referred to the article and the method; prerequisite: the guide), progression box added, page moved to MDX.

Revised on 5 October 2026: lab 4 duration aligned with its steps (1 h 35), bringing the total to about 13 h 20.