Skip to content

TechnicalPractitioner

Understanding generative AI observability

For: engineers and SREs · architectsPrerequisites: Good knowledge of classical observability (logs, metrics, traces) and the basics of LLM applications.

Reading mode

Scope: this guide covers generative AI observability, that is LLMs, RAG pipelines, agents and MCP tools; classic ML and computer vision are out of scope, see Beyond the LLM: GPUs and model quality.

Audience: SRE, platform and ML engineering teams who run, or are about to run, AI features in production. Prerequisites: working knowledge of classical observability (logs, metrics, traces) and of the basics of LLM applications. Goal: know what to measure, in what order, with which tools, and how to operate the platform over time.

An AI system can be up, fast and wrong. The guide starts there and works down into the detail: the signals to capture, maturity levels to break up the effort, a platform built in seven steps, evaluators you can actually trust. It also covers day-to-day operations, with two constraints that come up everywhere, budget and personal data.

  1. Up, fast and wrong. Classical observability covers system health (the horizontal axis). AI observability adds output quality (the vertical axis): is the answer correct, faithful and safe?
  2. Five properties break classical assumptions: non-determinism, natural-language payloads, quality as a runtime property, drift, and a variable cost per request (a single agent loop can burn the cost of thousands of requests).
  3. Five signals, one key. Logs, metrics and traces, plus evaluations and feedback. The trace ID links every span, every score and every feedback event into one queryable object.
  4. Every layer of the AI stack emits its own signals, from the application down to the model, the retriever and the tools. All of them correlate by trace ID.
  5. Four overlapping system types: LLM-only, RAG, agents, MCP tools. Real applications mix them, so the platform must handle all four.
  6. A trace tree tells the story of a request. Evaluations live in the same trace as the request they score; feedback attaches later, indexed by trace ID.
  7. A six-level maturity ladder. Place yourself honestly: without an explicit observability strategy, an AI feature stays at Level 0. In my view, Level 3 (online evaluation) is a realistic six-month objective.
  8. Seven ordered steps: define the questions, instrument, collect, store, evaluate, visualize and alert, close the loop. You instrument to answer questions you have already written down; instrumenting first and hoping the questions show up later does not work.
  9. The OpenTelemetry Collector is the integration point. It redacts, enriches, tail-samples, batches and exports. Order matters: redaction runs before batching.
  10. A self-hostable reference architecture where each layer is independently replaceable, with its alternatives: OpenInference, Collector, VictoriaMetrics, Phoenix, Tempo, VictoriaLogs, Ragas, Grafana, to be compared layer by layer with Prometheus, Mimir, Langfuse, Loki or managed services.
  11. Three evaluator types with very different costs. Rule-based, classifier-based, LLM-as-judge. For example: cheap checks on 100 percent of traces, small judges on 10 percent, large judges on 1 percent plus flagged traces. Always evaluate every error and every negative feedback.
  12. Online and offline evaluation are both required. Online evaluation catches drift in production; offline evaluation acts as the gate before each upgrade.
  13. Operating the platform: tail sampling by default, bounded metrics dimensions, high cardinality on traces only, redaction in the collector, RBAC and an audit log on trace queries, an egress allowlist for judges.
  14. Version, replay, experiment. Version every artifact that influences output, replay old traces against new versions, use shadow mode, A/B and canaries, and block the build when a score regresses.
  15. Avoid the anti-patterns and close the loop. Without a weekly review that turns failed traces into regression tests, the platform is an expensive read-only archive. Start from your current level (see the appendix).
  1. Foundations: the two axes, the five signals, the vocabulary.
  2. Scope and taxonomy of AI systems: LLM-only, RAG, agents, MCP, and the shape of their traces.
  3. The maturity roadmap: from Level 0 (blind operation) to Level 5 (closed loop).
  4. Implementation in seven steps: from the questions to the closed loop, with the OpenTelemetry GenAI attributes.
  5. Evaluator design: catalog, LLM-as-judge patterns, calibration, aggregation, sampling.
  6. The tools landscape: self-hosting or managed service, licenses, reference stacks.
  7. Operating the platform: sampling, cost, cardinality, PII, security, ownership.
  8. Versioning, replay and experimentation: safe iteration at scale.
  9. Anti-patterns and common pitfalls: twelve traps and their fixes.
  10. Appendix: glossary, references, next steps by level.

Several pieces on this site deal with generative AI observability. To avoid diverging versions, each topic has a single home; the other pieces summarize it in one sentence and link here.

TopicReference lesson
The five signals (logs, metrics, traces, evaluations, user feedback)1. Foundations
The four system types (LLM-only, RAG, agents, MCP)2. Scope and taxonomy
Maturity levels 0 to 53. The maturity roadmap
Evaluator design, LLM-as-judge, judge sampling, online and offline evaluation5. Evaluator design (and chapter 8 for replay and CI regression gating)
Tools landscape and component licenses6. The tools landscape
Cost of observability7. Operating the platform, section 7.2
Anti-patterns9. Anti-patterns and common pitfalls

Other topics live elsewhere and the guide links to them: the drift taxonomy (data, concept, pipeline) in module 3 of the labs; the OpenTelemetry GenAI conventions (versions, renaming history, content capture, MCP) in the production article; quality SLOs, the cost of non-observability, the AI Act and the RACI in the GenAI method.

Revised on 4 October 2026: title “Understanding generative AI observability”, explicit scope (LLM, RAG, agents, MCP), progression box, table of the topics this guide is the reference for, links to the labs, the production article and the method.