Skip to content

TechnicalPractitioner

Module 1: why an LLM cannot be observed like a classic application

For: engineers and SREs · architectsPrerequisites: Have already called an LLM API and know a classic observability tool.

Classic applicationLLM application
Behaviordeterministic: same input, same outputstatistical: same input, variable outputs
Errorsvisible (4xx, 5xx, exceptions)semantic: the model is politely wrong, in 200 OK
Infrastructure metricsoften enoughcan stay green while quality degrades
Quality guaranteeunit testscontinuous evaluation in production
Costtied to infrastructure (CPU, RAM, IO)per request, depending on context size

In practice, a classic APM dashboard can show perfect latency and zero errors while the assistant gives wrong answers to part of its users.

These three axes are not signals: they are the three questions you ask of an LLM system, and they must be read together. To answer them, you draw on the five signals defined in chapter 1 of the guide: logs, metrics, traces, evaluations and user feedback.

AxisQuestionTypical metricsSignals used
Qualityis the answer relevant, faithful to the context, correct?faithfulness, context recall, hallucination rate, user feedbackevaluations, user feedback; traces for diagnosis
Costhow much per request, per customer, per feature, per model?input and output tokens, euros per request, monthly cost, cache ratemetrics (tokens), traces (usage of each request)
Performancedoes the user get the answer on time and without error?time to first token (TTFT), p95 and p99 latency, error rate, throughputmetrics, traces (latency per step), logs (errors)

Looking at one axis alone is misleading. A drop in cost can come from a truncated context that degrades quality; a rise in latency can come from a better model that improves quality.

An LLM call goes through five steps. Each one produces spans, metrics and logs, and carries its own attributes.

Client Pre-processing Model Post-processing Response
app, API, --> prompt, --> LLM --> parsing, --> evaluation,
agent RAG retrieval inference guardrails log
tenant_id rag_stage model guardrail_action quality_score
feature top_k tokens parser_status hallucination
user_id docs_retrieved latency

Observing a call covers these five points. If only the model call is instrumented, the other four steps stay in the dark.

The labs only need to place their stack among three families of tools, plus an instrumentation layer common to all of them:

FamilyExamples
General-purpose SaaS observability platformsDatadog LLM Observability, New Relic AI Monitoring, Grafana Cloud
LLM-specialized platforms, SaaS or self-hostedLangfuse, Arize Phoenix or Arize AX, LangSmith, Helicone
Assembled self-hosted stackOpenTelemetry Collector, VictoriaMetrics, Prometheus or Mimir, Tempo or Jaeger, Grafana, Loki
Instrumentation layer (libraries, not storage products)OpenTelemetry SDK and instrumentations with the GenAI conventions, OpenLLMetry (Traceloop), OpenInference (Arize)

The full tools landscape, the selection criteria (data sovereignty, cost, integration) and the licenses are covered in chapter 6 of the guide. The versions and licenses of the kit components are in the overview; keep in mind that Phoenix is under Elastic License 2.0 (source-available, not open source).

These labs use a self-hosted stack, complemented by Phoenix, for a teaching reason: you see every link in the chain. The choice remains reversible, because OpenTelemetry instrumentation is portable: the same spans and metrics can be sent to any of the families above.

flowchart TB
  APP["LLM application<br/>Python, OTel SDK,<br/>GenAI conventions"] -->|"OTLP"| COL["OTel Collector<br/>reception, enrichment,<br/>routing"]
  COL -->|"metrics"| VM["VictoriaMetrics"]
  COL -->|"traces"| TP["Tempo"]
  COL -->|"traces"| PX["Phoenix<br/>exploration, evaluation"]
  VM --> GR["Grafana<br/>dashboards, alerts"]
  TP --> GR

The application is instrumented only once. The Collector then distributes the signals: metrics to VictoriaMetrics, the same traces to Tempo and to Phoenix. Grafana is the reading point for metrics, traces and alerts; Phoenix has its own interface to explore and evaluate conversations.

An LLM fails silently, with semantic rather than technical errors, which forces you to read quality, cost and performance together. Instrumentation covers the five steps of the chain, not only the model call. With OpenTelemetry at the center, the choice of backends remains reversible.

Revised on 2 October 2026: tool families are compared on the same criteria (getting started, cost, data location, operations), instrumentation libraries are separated from storage products, licenses are specified (Phoenix under Elastic License 2.0, source-available) and Phoenix now receives traces from the Collector.

Revised on 4 October 2026: the three axes are presented as questions and mapped onto the guide’s five signals; the “Choosing a stack” section is reduced to what the labs need and points to chapter 6 of the guide for the tools landscape and licenses.