Technical
Appendix: glossary, references, next steps
For: engineers and SREs · architectsPrerequisites: Know the maturity levels from chapter 3.
A.1 Glossary
Section titled “A.1 Glossary”Attribute: A key-value pair attached to a span, used for filtering and aggregation. Example: gen_ai.request.model = "provider-model".
Calibration: The process of measuring how an LLM-as-judge’s scores correlate with human judgment on a curated set. Required before any LLM-as-judge metric can be trusted.
Drift: A change over time that can degrade an AI system’s outputs. Three families: data (the inputs change: input drift or covariate shift, embedding drift, prior shift, virtual drift), concept (the reality the system describes changes: concept drift), pipeline (a component changes: model or provider silently changed, prompt template, RAG chunking). Output drift is not a fourth family: it is the symptom revealed by evaluations, whose cause belongs to one of the three. The taxonomy and its correspondence table are in module 3 of the labs, which is the reference.
Faithfulness: Whether an answer is grounded in the provided context. A faithful answer may still be wrong if the context is wrong. Distinct from factual correctness.
Gold set: A curated dataset of inputs with known ideal outputs, used for offline evaluation.
Grounding: The property that an answer is supported by a verifiable source rather than parametric knowledge.
Hallucination: Output that asserts facts not supported by training data, provided context or any verifiable source.
Head sampling: Sampling decision taken at trace start, by hash or probability. Stateless and cheap, but cannot select on outcome.
LLM-as-judge: Evaluation method where a model scores another model’s output against a rubric. A common approach for semantic quality, to be calibrated against human judgment.
OpenInference: Semantic conventions and instrumentation libraries originally from Arize, designed for AI tracing. They export over OTLP, but with their own attributes (openinference.span.kind, llm.*) rather than the OpenTelemetry gen_ai.* attributes: you either convert attributes (Collector transform processor) or pick a single convention.
OpenTelemetry GenAI: Semantic conventions extending OpenTelemetry to cover language model operations. Still in development status, maintained in a dedicated repository since June 2026.
Regression set: A dataset of inputs that previously failed, used to gate future changes. Grows over time as new failures are triaged.
Span: A single timed operation, such as one LLM call, one retrieval, one tool invocation. Spans nest to form trace trees.
Tail sampling: Sampling decision taken at trace end, by policy. Requires buffering complete traces but allows selection on outcome.
Trace: The full execution path of a request, composed of spans linked by parent IDs. The fundamental unit of AI observability.
A.2 Reference material
Section titled “A.2 Reference material”- First version of this guide, “Understanding AI Observability” (May 2026): forge.erythix.tech/labstraining/training/AIOBS
- OpenTelemetry GenAI semantic conventions (development status, dedicated repository since semantic-conventions v1.42.0, June 2026): github.com/open-telemetry/semantic-conventions-genai
- OpenInference specification: github.com/Arize-ai/openinference
- Phoenix documentation: arize.com/docs/phoenix
- Langfuse documentation (acquired by ClickHouse in January 2026): langfuse.com/docs
- Ragas documentation: docs.ragas.io
- DeepEval documentation: deepeval.com/docs/getting-started
- OpenLLMetry (Traceloop, acquired by ServiceNow in March 2026): github.com/traceloop/openllmetry
- VictoriaMetrics documentation: docs.victoriametrics.com
- OpenTelemetry Collector tail sampling processor: github.com/open-telemetry/opentelemetry-collector-contrib/tree/main/processor/tailsamplingprocessor
A.3 Recommended next steps by starting level
Section titled “A.3 Recommended next steps by starting level”If you are at Level 0
Section titled “If you are at Level 0”Target Level 1 in two weeks. Add structured logging around every LLM call. Capture prompt, response, model, token counts, latency, error and a request ID. Aggregate token usage into a basic cost dashboard. Do not start with OpenTelemetry, do not start with evaluation: ship the logs first.
If you are at Level 1
Section titled “If you are at Level 1”Target Level 2 in one to two months. Adopt OpenTelemetry GenAI conventions and migrate your custom logging to span attributes. Deploy an OpenTelemetry Collector and a trace backend. If you want an AI-specific UI without extra integration, a self-hosted tool such as Phoenix or Langfuse is a good starting point; if you already run Tempo or Jaeger, start with it. Validate that you can answer the questions from Step 1 of Chapter 4.
If you are at Level 2
Section titled “If you are at Level 2”Target Level 3 in two to three months. Add online evaluation on the most critical path. Start with three evaluators: format compliance, toxicity and one domain-specific scorer. Wire the scores to alerts. Establish a weekly review of failed and low-score traces.
If you are at Level 3
Section titled “If you are at Level 3”Target Level 4 in three months. Add embedding drift monitoring. Wire user feedback signals into the trace store. Build a maintained gold set for RAG evaluation. Begin offline evaluation in CI to gate model and prompt changes.
If you are at Level 4
Section titled “If you are at Level 4”Target Level 5 within a quarter. Formalize the regression dataset, the few-shot example set, and the upgrade gate. Automate re-indexing on drift. Treat the observability platform as the source of truth for model and prompt iteration.
Back to the overview.
Revised on 2 October 2026: GenAI conventions link moved to the semantic-conventions-genai repository, Phoenix and DeepEval links updated, Langfuse and Traceloop acquisitions, OpenInference’s own attributes clarified, first version of the guide cited.
Revised on 4 October 2026: drift definition aligned on the site’s three families (data, concept, pipeline), with a link to module 3 of the labs.