Skip to content

TechnicalExpert

Part I. The framework

For: engineers and SREs · architectsPrerequisites: Basic notions of metrics, traces and OpenTelemetry.

Classical observability was designed for loud failures: a 500 code, a timeout, an exception. Systems built on LLMs fail differently. Their three dominant pathologies are silent:

  • quality degrades with no error raised: hallucination, incomplete answer, slow drift.
  • cost drifts with no alert: prompt bloat, agent loops.
  • security is compromised at the semantic level, not the network level: tool poisoning, injection.

Statement of the model: for any signal you do not measure, the relevant question is not “is the service down?” but “how long does the failure accumulate before you see it?”. The value of a GenAI observability setup is measured by its ability to make the silent visible. Everything else in the method follows from this, and the impact analysis in Part V quantifies it.

flowchart TB
  subgraph D1["Dimension 1: the 4 layers (where the signal occurs)"]
    L1["Infrastructure: GPU, network, vector store, MCP server"]
    L2["Model / inference: latency, tokens, saturation"]
    L3["Application: RAG, prompt, orchestration, guardrails"]
    L4["Agent / task: trajectory, tools, end-to-end success"]
    L1 --> L2 --> L3 --> L4
  end
  subgraph D2["Dimension 2: the 4 axes (why you monitor)"]
    A1["Reliability"]
    A2["Availability"]
    A3["Cost"]
    A4["Compliance"]
  end
  subgraph D3["Dimension 3: the 5 questions (per signal)"]
    Q["What -> Why -> How -> Threshold -> Impact if absent"]
  end
  subgraph D4["Dimension 4: maturity (the guide's levels 0 to 5)"]
    M["0. Blind -> 1. Basic telemetry -> 2. Tracing -> 3. Online evaluation -> 4. Drift and feedback -> 5. Closed loop"]
  end
  D1 --> D2 --> D3 --> D4

Every signal belongs to a layer. Structuring rule: the cascade. A low failure surfaces as a high symptom. Without correlation through a single trace_id across layers, you observe symptoms with no cause.

flowchart LR
  C2["Layer 2: provider rate-limit at 10:47"] --> C3["Layer 3: retries, prompt latency"] --> C4["Layer 4: task success rate dropping"]
  C4 -.->|"downward diagnosis via trace_id"| C2

The site describes the stack of a generative AI application in five layers in AI in cross-section: data, infrastructure, model, orchestration, application. The method’s four layers map onto them as follows (mind the word “application”, which does not mean the same thing):

Method layerAI in cross-section layer(s)
1. Infrastructure: GPU, network, vector store, MCP serverinfrastructure (GPUs, network, servers); the indexed content of the vector store belongs to data; the tools exposed by an MCP server are called from orchestration
2. Model / inference: latency, tokens, saturationmodel, at its boundary with infrastructure: TTFT, tokens per second and saturation are measured at the inference server, which the diagram places in infrastructure
3. Application: RAG, prompt, orchestration, guardrailsorchestration (prompts, document retrieval); output guardrails also touch the application layer (policy violations, personal data in outputs)
4. Agent / task: trajectory, tools, end-to-end successorchestration for the trajectory and tool calls; application for the success the user perceives
no dedicated layerdata: source freshness and coverage, input distribution, which the method handles in the RAG grid of Part III

Reliability, availability, cost, compliance. Each component is read along these four axes in the grids of Part III.

The template applied to every signal. You never instrument without answering “why” and “impact if absent”.

QuestionWhat it establishes
WhatThe precise signal: metric, trace, score, event
WhyThe decision it informs. A signal with no decision is noise
HowThe source and the instrument: OTel attribute, exporter, evaluator
ThresholdThe SLI and the SLO, hence the alert trigger
Impact if absentThe failure that accumulates without this signal, and its severity

The method used to have its own five-level scale (Logging, Tracing, Evaluation, Monitoring, Closed loop). The site keeps only one, that of the guide, chapter 3, with six levels from 0 to 5. Correspondence:

Guide levelWhat characterizes itFormer method level
0. Blind operationHTTP metrics only; cost discovered on the bill, quality through complaintsnone: the method’s scale started at logging
1. Basic telemetrystructured per-request logs (prompt, response, model, tokens, latency), aggregated cost, errors as metrics1. Logging
2. Structured tracingOpenTelemetry GenAI conventions, one span per logical step, queryable attributes2. Tracing
3. Online evaluationautomated evaluations attached to traces and exposed as metrics, alertable quality regressions3. Evaluation
4. Drift and feedback integrationembedding drift, batch RAG quality, user feedback correlated with traces4. Monitoring (approximate match: the method did not detail this level)
5. Closed loopfailing traces become regression cases, changes gated on the regression set5. Closed loop

The quality SLOs of Part IV assume at least level 3: without online evaluation there is no sample to measure.

In my view, many organizations stop at level 2 and believe they observe their system. They see what happened, not whether it was good. The value jump is at level 3, mastery at level 5.

The signals are those of the guide, chapter 1: logs, metrics, traces, evaluations, user feedback, correlated by the trace ID. The method adds the typical storage of its reference configuration:

SignalTypical storage
LogsVictoriaLogs, Loki
MetricsVictoriaMetrics, Prometheus
TracesOTLP backend (Tempo, Jaeger)
EvaluationsLangfuse, Arize Phoenix
User feedbackattached to the trace by its ID, next to the evaluation scores

The “content events” (prompt, tool call, result) listed in the first version of the method are not a separate signal: they belong to logs and traces, as content capture (rule below).

The instrumentation baseline to adopt: the OpenTelemetry GenAI semantic conventions, in Development status. Their versions, their history (including the move to a dedicated repository since 1.42), their attributes and content capture are detailed only once on the site, in the article Observing an LLM system, §2. The method relies mostly on gen_ai.request.model and gen_ai.response.model, gen_ai.provider.name, gen_ai.usage.input_tokens and gen_ai.usage.output_tokens (the basis of cost), gen_ai.response.finish_reasons (a first-order reliability signal) and gen_ai.operation.name. Its recording rules (Part II) start from the metrics of versions 1.40 and 1.41 (gen_ai.client.operation.duration, gen_ai.client.token.usage): see the revision note at the bottom of the page.

Content capture rule, aligned with the specification: instrumentations do not capture instructions, inputs and outputs by default; capturing them is an explicit opt-in. Three usage patterns: record nothing (the default); record the content on spans (gen_ai.system_instructions, gen_ai.input.messages, gen_ai.output.messages), reserved for cases where volume stays manageable and storage complies with regulation, for example in pre-production; store the content externally and record a reference on the span, the pattern the specification recommends in production. In a regulated environment the method retains the third, with masking at the collector. Content never goes into a metric label (see the guide, 9.6).


Revised on 2 October 2026: since version 1.42 (June 2026), the GenAI conventions live in the semantic-conventions-genai repository. At that date, its main branch, whose changes are not yet published in a release, lists gen_ai.client.operation.duration as recommended rather than required, and replaces gen_ai.client.token.usage with gen_ai.client.inference.usage.* counters (input, output, cache and reasoning tokens) and gen_ai.client.inference.operation.* histograms. Check which version your instrumentation emits before writing your rules.

Revised on 4 October 2026: signals aligned on the guide’s five (content events belong to logs and traces), correspondence table to the five layers of AI in cross-section, maturity scale replaced by the guide’s levels 0 to 5, OpenTelemetry conventions recap reduced to a summary with a pointer to the article, content capture rule aligned with the specification.