Technical
Module 1: why an LLM cannot be observed like a classic application
For: engineers and SREs · architectsPrerequisites: Have already called an LLM API and know a classic observability tool.
Two opposite behaviors
Section titled “Two opposite behaviors”| Classic application | LLM application | |
|---|---|---|
| Behavior | deterministic: same input, same output | statistical: same input, variable outputs |
| Errors | visible (4xx, 5xx, exceptions) | semantic: the model is politely wrong, in 200 OK |
| Infrastructure metrics | often enough | can stay green while quality degrades |
| Quality guarantee | unit tests | continuous evaluation in production |
| Cost | tied to infrastructure (CPU, RAM, IO) | per request, depending on context size |
In practice, a classic APM dashboard can show perfect latency and zero errors while the assistant gives wrong answers to part of its users.
Three observation axes to industrialize
Section titled “Three observation axes to industrialize”These three axes are not signals: they are the three questions you ask of an LLM system, and they must be read together. To answer them, you draw on the five signals defined in chapter 1 of the guide: logs, metrics, traces, evaluations and user feedback.
| Axis | Question | Typical metrics | Signals used |
|---|---|---|---|
| Quality | is the answer relevant, faithful to the context, correct? | faithfulness, context recall, hallucination rate, user feedback | evaluations, user feedback; traces for diagnosis |
| Cost | how much per request, per customer, per feature, per model? | input and output tokens, euros per request, monthly cost, cache rate | metrics (tokens), traces (usage of each request) |
| Performance | does the user get the answer on time and without error? | time to first token (TTFT), p95 and p99 latency, error rate, throughput | metrics, traces (latency per step), logs (errors) |
Looking at one axis alone is misleading. A drop in cost can come from a truncated context that degrades quality; a rise in latency can come from a better model that improves quality.
Where to instrument
Section titled “Where to instrument”An LLM call goes through five steps. Each one produces spans, metrics and logs, and carries its own attributes.
Client Pre-processing Model Post-processing Responseapp, API, --> prompt, --> LLM --> parsing, --> evaluation,agent RAG retrieval inference guardrails log
tenant_id rag_stage model guardrail_action quality_scorefeature top_k tokens parser_status hallucinationuser_id docs_retrieved latencyObserving a call covers these five points. If only the model call is instrumented, the other four steps stay in the dark.
Choosing a stack
Section titled “Choosing a stack”The labs only need to place their stack among three families of tools, plus an instrumentation layer common to all of them:
| Family | Examples |
|---|---|
| General-purpose SaaS observability platforms | Datadog LLM Observability, New Relic AI Monitoring, Grafana Cloud |
| LLM-specialized platforms, SaaS or self-hosted | Langfuse, Arize Phoenix or Arize AX, LangSmith, Helicone |
| Assembled self-hosted stack | OpenTelemetry Collector, VictoriaMetrics, Prometheus or Mimir, Tempo or Jaeger, Grafana, Loki |
| Instrumentation layer (libraries, not storage products) | OpenTelemetry SDK and instrumentations with the GenAI conventions, OpenLLMetry (Traceloop), OpenInference (Arize) |
The full tools landscape, the selection criteria (data sovereignty, cost, integration) and the licenses are covered in chapter 6 of the guide. The versions and licenses of the kit components are in the overview; keep in mind that Phoenix is under Elastic License 2.0 (source-available, not open source).
These labs use a self-hosted stack, complemented by Phoenix, for a teaching reason: you see every link in the chain. The choice remains reversible, because OpenTelemetry instrumentation is portable: the same spans and metrics can be sent to any of the families above.
The target architecture
Section titled “The target architecture”flowchart TB APP["LLM application<br/>Python, OTel SDK,<br/>GenAI conventions"] -->|"OTLP"| COL["OTel Collector<br/>reception, enrichment,<br/>routing"] COL -->|"metrics"| VM["VictoriaMetrics"] COL -->|"traces"| TP["Tempo"] COL -->|"traces"| PX["Phoenix<br/>exploration, evaluation"] VM --> GR["Grafana<br/>dashboards, alerts"] TP --> GR
The application is instrumented only once. The Collector then distributes the signals: metrics to VictoriaMetrics, the same traces to Tempo and to Phoenix. Grafana is the reading point for metrics, traces and alerts; Phoenix has its own interface to explore and evaluate conversations.
In summary
Section titled “In summary”An LLM fails silently, with semantic rather than technical errors, which forces you to read quality, cost and performance together. Instrumentation covers the five steps of the chain, not only the model call. With OpenTelemetry at the center, the choice of backends remains reversible.
Revised on 2 October 2026: tool families are compared on the same criteria (getting started, cost, data location, operations), instrumentation libraries are separated from storage products, licenses are specified (Phoenix under Elastic License 2.0, source-available) and Phoenix now receives traces from the Collector.
Revised on 4 October 2026: the three axes are presented as questions and mapped onto the guide’s five signals; the “Choosing a stack” section is reduced to what the labs need and points to chapter 6 of the guide for the tools landscape and licenses.