Skip to content

TechnicalPractitioner

Appendix: glossary, references, next steps

For: engineers and SREs · architectsPrerequisites: Know the maturity levels from chapter 3.

Attribute: A key-value pair attached to a span, used for filtering and aggregation. Example: gen_ai.request.model = "provider-model".

Calibration: The process of measuring how an LLM-as-judge’s scores correlate with human judgment on a curated set. Required before any LLM-as-judge metric can be trusted.

Drift: A change over time that can degrade an AI system’s outputs. Three families: data (the inputs change: input drift or covariate shift, embedding drift, prior shift, virtual drift), concept (the reality the system describes changes: concept drift), pipeline (a component changes: model or provider silently changed, prompt template, RAG chunking). Output drift is not a fourth family: it is the symptom revealed by evaluations, whose cause belongs to one of the three. The taxonomy and its correspondence table are in module 3 of the labs, which is the reference.

Faithfulness: Whether an answer is grounded in the provided context. A faithful answer may still be wrong if the context is wrong. Distinct from factual correctness.

Gold set: A curated dataset of inputs with known ideal outputs, used for offline evaluation.

Grounding: The property that an answer is supported by a verifiable source rather than parametric knowledge.

Hallucination: Output that asserts facts not supported by training data, provided context or any verifiable source.

Head sampling: Sampling decision taken at trace start, by hash or probability. Stateless and cheap, but cannot select on outcome.

LLM-as-judge: Evaluation method where a model scores another model’s output against a rubric. A common approach for semantic quality, to be calibrated against human judgment.

OpenInference: Semantic conventions and instrumentation libraries originally from Arize, designed for AI tracing. They export over OTLP, but with their own attributes (openinference.span.kind, llm.*) rather than the OpenTelemetry gen_ai.* attributes: you either convert attributes (Collector transform processor) or pick a single convention.

OpenTelemetry GenAI: Semantic conventions extending OpenTelemetry to cover language model operations. Still in development status, maintained in a dedicated repository since June 2026.

Regression set: A dataset of inputs that previously failed, used to gate future changes. Grows over time as new failures are triaged.

Span: A single timed operation, such as one LLM call, one retrieval, one tool invocation. Spans nest to form trace trees.

Tail sampling: Sampling decision taken at trace end, by policy. Requires buffering complete traces but allows selection on outcome.

Trace: The full execution path of a request, composed of spans linked by parent IDs. The fundamental unit of AI observability.

Section titled “A.3 Recommended next steps by starting level”

Target Level 1 in two weeks. Add structured logging around every LLM call. Capture prompt, response, model, token counts, latency, error and a request ID. Aggregate token usage into a basic cost dashboard. Do not start with OpenTelemetry, do not start with evaluation: ship the logs first.

Target Level 2 in one to two months. Adopt OpenTelemetry GenAI conventions and migrate your custom logging to span attributes. Deploy an OpenTelemetry Collector and a trace backend. If you want an AI-specific UI without extra integration, a self-hosted tool such as Phoenix or Langfuse is a good starting point; if you already run Tempo or Jaeger, start with it. Validate that you can answer the questions from Step 1 of Chapter 4.

Target Level 3 in two to three months. Add online evaluation on the most critical path. Start with three evaluators: format compliance, toxicity and one domain-specific scorer. Wire the scores to alerts. Establish a weekly review of failed and low-score traces.

Target Level 4 in three months. Add embedding drift monitoring. Wire user feedback signals into the trace store. Build a maintained gold set for RAG evaluation. Begin offline evaluation in CI to gate model and prompt changes.

Target Level 5 within a quarter. Formalize the regression dataset, the few-shot example set, and the upgrade gate. Automate re-indexing on drift. Treat the observability platform as the source of truth for model and prompt iteration.

Back to the overview.

Revised on 2 October 2026: GenAI conventions link moved to the semantic-conventions-genai repository, Phoenix and DeepEval links updated, Langfuse and Traceloop acquisitions, OpenInference’s own attributes clarified, first version of the guide cited.

Revised on 4 October 2026: drift definition aligned on the site’s three families (data, concept, pipeline), with a link to module 3 of the labs.