Technical
Understanding generative AI observability
For: engineers and SREs · architectsPrerequisites: Good knowledge of classical observability (logs, metrics, traces) and the basics of LLM applications.
Scope: this guide covers generative AI observability, that is LLMs, RAG pipelines, agents and MCP tools; classic ML and computer vision are out of scope, see Beyond the LLM: GPUs and model quality.
Audience: SRE, platform and ML engineering teams who run, or are about to run, AI features in production. Prerequisites: working knowledge of classical observability (logs, metrics, traces) and of the basics of LLM applications. Goal: know what to measure, in what order, with which tools, and how to operate the platform over time.
An AI system can be up, fast and wrong. The guide starts there and works down into the detail: the signals to capture, maturity levels to break up the effort, a platform built in seven steps, evaluators you can actually trust. It also covers day-to-day operations, with two constraints that come up everywhere, budget and personal data.
The guide in fifteen points
Section titled “The guide in fifteen points”- Up, fast and wrong. Classical observability covers system health (the horizontal axis). AI observability adds output quality (the vertical axis): is the answer correct, faithful and safe?
- Five properties break classical assumptions: non-determinism, natural-language payloads, quality as a runtime property, drift, and a variable cost per request (a single agent loop can burn the cost of thousands of requests).
- Five signals, one key. Logs, metrics and traces, plus evaluations and feedback. The trace ID links every span, every score and every feedback event into one queryable object.
- Every layer of the AI stack emits its own signals, from the application down to the model, the retriever and the tools. All of them correlate by trace ID.
- Four overlapping system types: LLM-only, RAG, agents, MCP tools. Real applications mix them, so the platform must handle all four.
- A trace tree tells the story of a request. Evaluations live in the same trace as the request they score; feedback attaches later, indexed by trace ID.
- A six-level maturity ladder. Place yourself honestly: without an explicit observability strategy, an AI feature stays at Level 0. In my view, Level 3 (online evaluation) is a realistic six-month objective.
- Seven ordered steps: define the questions, instrument, collect, store, evaluate, visualize and alert, close the loop. You instrument to answer questions you have already written down; instrumenting first and hoping the questions show up later does not work.
- The OpenTelemetry Collector is the integration point. It redacts, enriches, tail-samples, batches and exports. Order matters: redaction runs before batching.
- A self-hostable reference architecture where each layer is independently replaceable, with its alternatives: OpenInference, Collector, VictoriaMetrics, Phoenix, Tempo, VictoriaLogs, Ragas, Grafana, to be compared layer by layer with Prometheus, Mimir, Langfuse, Loki or managed services.
- Three evaluator types with very different costs. Rule-based, classifier-based, LLM-as-judge. For example: cheap checks on 100 percent of traces, small judges on 10 percent, large judges on 1 percent plus flagged traces. Always evaluate every error and every negative feedback.
- Online and offline evaluation are both required. Online evaluation catches drift in production; offline evaluation acts as the gate before each upgrade.
- Operating the platform: tail sampling by default, bounded metrics dimensions, high cardinality on traces only, redaction in the collector, RBAC and an audit log on trace queries, an egress allowlist for judges.
- Version, replay, experiment. Version every artifact that influences output, replay old traces against new versions, use shadow mode, A/B and canaries, and block the build when a score regresses.
- Avoid the anti-patterns and close the loop. Without a weekly review that turns failed traces into regression tests, the platform is an expensive read-only archive. Start from your current level (see the appendix).
Contents
Section titled “Contents”- Foundations: the two axes, the five signals, the vocabulary.
- Scope and taxonomy of AI systems: LLM-only, RAG, agents, MCP, and the shape of their traces.
- The maturity roadmap: from Level 0 (blind operation) to Level 5 (closed loop).
- Implementation in seven steps: from the questions to the closed loop, with the OpenTelemetry GenAI attributes.
- Evaluator design: catalog, LLM-as-judge patterns, calibration, aggregation, sampling.
- The tools landscape: self-hosting or managed service, licenses, reference stacks.
- Operating the platform: sampling, cost, cardinality, PII, security, ownership.
- Versioning, replay and experimentation: safe iteration at scale.
- Anti-patterns and common pitfalls: twelve traps and their fixes.
- Appendix: glossary, references, next steps by level.
Topics this guide is the reference for
Section titled “Topics this guide is the reference for”Several pieces on this site deal with generative AI observability. To avoid diverging versions, each topic has a single home; the other pieces summarize it in one sentence and link here.
| Topic | Reference lesson |
|---|---|
| The five signals (logs, metrics, traces, evaluations, user feedback) | 1. Foundations |
| The four system types (LLM-only, RAG, agents, MCP) | 2. Scope and taxonomy |
| Maturity levels 0 to 5 | 3. The maturity roadmap |
| Evaluator design, LLM-as-judge, judge sampling, online and offline evaluation | 5. Evaluator design (and chapter 8 for replay and CI regression gating) |
| Tools landscape and component licenses | 6. The tools landscape |
| Cost of observability | 7. Operating the platform, section 7.2 |
| Anti-patterns | 9. Anti-patterns and common pitfalls |
Other topics live elsewhere and the guide links to them: the drift taxonomy (data, concept, pipeline) in module 3 of the labs; the OpenTelemetry GenAI conventions (versions, renaming history, content capture, MCP) in the production article; quality SLOs, the cost of non-observability, the AI Act and the RACI in the GenAI method.
What next
Section titled “What next”- Practice: LLM observability: the labs puts chapters 4 and 5 to work on a self-hosted stack (instrumentation, RAG, drift, cost).
- Go to production: the article Observing an LLM system gives the current OpenTelemetry GenAI conventions and the deployment playbooks (Collector, sampling, RAG, MCP, privacy), followed by VictoriaMetrics as an LLM observability backend.
- Master: the GenAI method covers quality SLOs, quantified impact, compliance and organization.
- Go beyond: Beyond the LLM: GPUs and model quality for GPU infrastructure, classic ML and statistical drift.
Revised on 4 October 2026: title “Understanding generative AI observability”, explicit scope (LLM, RAG, agents, MCP), progression box, table of the topics this guide is the reference for, links to the labs, the production article and the method.