Skip to content

TechnicalPractitioner

2. Scope and taxonomy of AI systems

For: engineers and SREs · architectsPrerequisites: Have read chapter 1 of the guide.

AI observability covers four overlapping system types. Each has distinct instrumentation requirements and characteristic failure modes. The same observability stack must accommodate all four because real applications mix them freely. A production agent typically performs retrieval and invokes MCP tools: three of the four types are exercised by a single request.

This four-type taxonomy is the reference for every piece on this site about generative AI observability: the labs, the production article and the GenAI method build on it. Classic ML and computer vision systems are out of this scope: see Beyond the LLM: GPUs and model quality.

The four system types: LLM-only, RAG, agentic, MCP and tools, with their typical failure modes

Figure 3. The four system types. Failure modes are illustrative, not exhaustive.

A single call to a language model with no retrieval and no tool use. The simplest case, and often the first one shipped to production.

  • Prompt, including the system prompt and the user prompt as separate fields.
  • Response, with finish_reason (stop, length, content_filter, tool_call, error in the OpenTelemetry conventions vocabulary).
  • Token usage, both input and output, distinguishing cached tokens where the provider exposes them.
  • Latency, end to end and time to first token if streaming.
  • Model identity, including the actually-served version (often differs from the requested name).
  • Sampling parameters: temperature, top-p, top-k, seed if set.
  • Hallucinated entities, dates or quotations.
  • Format violations: expected JSON, got prose.
  • Prompt injection through user content.
  • Runaway cost when streaming is uncapped and output exceeds expectations.
  • Provider-side outages and silent model swaps.

Retrieval-augmented generation chains a retriever (vector database, BM25 or hybrid) with a language model. Observability must additionally capture the retrieval stage and the relationship between retrieved context and generated answer. The retriever introduces a second axis of failure: the model can be perfect and still produce a wrong answer because the wrong documents were retrieved.

  • Retrieval query, which is often a rewrite of the user query.
  • Retrieved chunks, with their identifiers, source documents and similarity scores.
  • Top-k cutoff and any filters applied (date, source, tenant).
  • Reranking step: model identity, input chunks, output order.
  • Chunk identifiers in the final prompt, so each chunk can be tied to the answer that used it.
  • Context precision: fraction of retrieved chunks that are relevant to the query.
  • Context recall: fraction of relevant chunks that were retrieved.
  • Faithfulness: whether the answer is supported by the retrieved context, independent of whether that context is correct.
  • Answer relevancy: whether the answer addresses the question asked.
  • Low recall: the relevant document was never retrieved. The model invents a plausible answer.
  • Irrelevant top-k: similarity score was high but the chunks do not actually answer the query.
  • Contradiction: the answer contradicts the retrieved context (faithfulness failure).
  • Stale chunks: the index has not been updated since a relevant change.
  • Duplicate chunks: the same content appears multiple times, biasing the answer.

An agent iterates over a loop of plan, act, observe. Observability must capture the loop structure and the reasoning at each step. Agents are the hardest system type to observe because failure modes include unbounded behavior: loops, cost explosion, premature termination.

  • Planner reasoning at each step, when exposed by the model. For reasoning models, the chain of thought is itself a span attribute.
  • Tool calls with name, arguments and response payload.
  • Iteration count and the reason for loop termination (success, budget, max iterations, error).
  • Per-step cost and latency, with running totals over the trace.
  • Detection of repeated tool calls with identical arguments, which usually indicates a loop.
  • Infinite or near-infinite loops, bounded only by max-iterations or cost budget.
  • Wrong tool selection when several tools have overlapping descriptions.
  • Malformed tool arguments (wrong types, missing required fields, invented enums).
  • Premature termination when the agent answers without completing the task.
  • Cost explosion when context grows linearly with iterations.

The Model Context Protocol standardizes how language models invoke tools across providers. When MCP servers are part of the system, observability must capture the protocol level, not just the function call. The MCP server can become a shared bottleneck across multiple agents and applications.

  • MCP server identity and version.
  • Tool name, arguments and response payload.
  • Latency and error rate per tool, per server.
  • Authentication scope used for each call (the token or identity propagated).
  • Tool discovery events: when the agent first queried for the available tool list.
  • Cross-server tool selection when multiple MCP servers are connected.
  • Tool unavailability: the MCP server is down or slow.
  • Schema drift between server and client when the server publishes a new tool signature.
  • Scope escalation when an agent uses a tool requiring more permissions than the user has.
  • Slow tools blocking the agent loop: a single tool with a 30-second timeout can dominate trace cost.
  • Tool poisoning: the MCP server returns crafted output designed to manipulate the agent.

The same span library produces very different trace trees depending on the system type. Figure 4 shows a realistic RAG-agent trace combining all four types: a root span, an agent loop, two LLM calls, two retrieval steps, an MCP tool call and three evaluation spans attached after generation.

Trace tree of a RAG-agent request: root span, agent loop, LLM calls, retrievals, MCP tool call and evaluation spans

Figure 4. Trace tree of a single RAG-agent request. The same trace contains the four system types: LLM calls, retrieval, agent loop and MCP tool. Evaluations attach as child spans of the trace.

Three patterns to internalize from this shape:

  • Evaluations live in the same trace as the request they evaluate. They are not separate logs.
  • Async feedback (the user’s thumbs-down, for example) is attached out of band but indexed by trace ID.
  • Per-step latency is read off the bar widths. The largest bar is the dominant cost: optimize that first.

Next: 3. The maturity roadmap.

Revised on 4 October 2026: four-type taxonomy presented as the site’s reference, scope (classic ML and vision excluded) restated.