Technical
1. Foundations
For: engineers and SREs · architects · business and productPrerequisites: A rough idea of what logs, metrics and traces are.
1.1 What AI observability is
Section titled “1.1 What AI observability is”AI observability is the discipline of measuring, recording and interpreting the behavior of AI systems in production. It extends classical observability with signals specific to non-deterministic, semantic-content systems: prompts, completions, embeddings, retrieval contexts, tool calls and quality evaluations.
The core question is no longer whether the system is up and responding within latency budgets. The system can be up, fast and wrong. AI observability adds a second axis: is the output correct, faithful and safe?

Figure 1. The two axes of AI observability. Classical observability covers the horizontal axis. AI observability adds the vertical axis.
If your monitoring can tell you the system is up but cannot tell you the answers are right, you are operating in the bottom-right quadrant. That is the failure mode AI observability exists to prevent.
1.2 Why classical observability is insufficient
Section titled “1.2 Why classical observability is insufficient”Five structural properties of AI systems break the assumptions that underpin classical observability.
- Non-determinism: identical inputs produce different outputs across runs. Repeatability cannot be assumed.
- Semantic payloads: requests and responses are natural language, not structured data with stable schemas. The shape of the payload is the answer.
- Quality is a runtime property: there is no compile-time guarantee of correctness. A system that returned 100 perfect answers can return a hallucination on the 101st with no code change.
- Drift: model providers update weights, embeddings shift as documents change, data distributions evolve as user behavior changes. The same code can produce different outputs week to week.
- Variable cost per request: token counts depend on input and output length, with no predictable upper bound. A single agent loop can burn the cost of thousands of normal requests.
A 200 OK response from an LLM endpoint tells you nothing about whether the answer was correct, hallucinated, leaking PII or biased. Classical observability cannot answer the questions that matter for an AI feature in production.
1.3 The five signals of AI observability
Section titled “1.3 The five signals of AI observability”Traditional observability has three signals: logs, metrics, traces. AI observability adds two more: evaluations and feedback. These two close the loop between what the system did and whether it was useful.

Figure 2. The five signals. Classical observability supplies the first three. Evaluations and feedback are the AI-specific additions and are the defining capability of the discipline.
All five signals are correlated by a single identifier, the trace ID. Every span, every score, every feedback event belongs to one trace. This is what makes the platform queryable end to end. Without the trace ID as a common key, the signals devolve into disconnected silos.
1.4 Vocabulary
Section titled “1.4 Vocabulary”A handful of terms recur throughout this guide and across the tooling landscape. The definitions are not negotiable across vendors: use them precisely.
- Trace: the full execution path of a request, composed of spans linked by parent IDs.
- Span: a single timed operation, such as one LLM call, one retrieval, one tool invocation.
- Attribute: a key-value pair attached to a span, used for filtering and aggregation.
- Online evaluation: scoring that runs against live production traces, usually asynchronously.
- Offline evaluation: scoring that runs against a fixed dataset, typically in CI or as a scheduled batch.
- LLM-as-judge: evaluation method where a model scores another model’s output against a rubric.
- Drift: a change over time in the inputs, in the reality the system describes or in the pipeline itself, which can degrade outputs. Three families: data, concept, pipeline (see the glossary).
- Grounding: the property that an answer is supported by a verifiable source rather than parametric knowledge.
- Hallucination: output that asserts facts not supported by training data, provided context or any verifiable source.
- Faithfulness: the property of being grounded in the provided context, scored as a continuous quality.
- Gold set: a curated dataset of inputs with known ideal outputs, used for offline evaluation.
- Regression set: a dataset of inputs that previously failed, used to gate future changes.
1.5 What AI observability is not
Section titled “1.5 What AI observability is not”Three frequent confusions worth removing early.
- AI observability is not MLOps. MLOps covers model training, packaging and deployment. AI observability covers what happens after deployment.
- AI observability is not just expensive logging. Capturing prompts in logs is the starting point, not the end state. Without evaluation, drift detection and feedback correlation, the data accumulates and nothing improves.
- AI observability is not a single product. It is a stack assembled from instrumentation, collection, storage, evaluation and visualization. No single tool covers all five competently across all four system types.
Next: 2. Scope and taxonomy of AI systems.
Revised on 4 October 2026: the five signals presented as the site’s reference model, with how the labs’ axes, the method’s content events and the article’s planes map onto them.