Technical
Part I. The framework
For: engineers and SREs · architectsPrerequisites: Basic notions of metrics, traces and OpenTelemetry.
1. The silent failure model
Section titled “1. The silent failure model”Classical observability was designed for loud failures: a 500 code, a timeout, an exception. Systems built on LLMs fail differently. Their three dominant pathologies are silent:
- quality degrades with no error raised: hallucination, incomplete answer, slow drift.
- cost drifts with no alert: prompt bloat, agent loops.
- security is compromised at the semantic level, not the network level: tool poisoning, injection.
Statement of the model: for any signal you do not measure, the relevant question is not “is the service down?” but “how long does the failure accumulate before you see it?”. The value of a GenAI observability setup is measured by its ability to make the silent visible. Everything else in the method follows from this, and the impact analysis in Part V quantifies it.
2. The four-dimension framework
Section titled “2. The four-dimension framework”flowchart TB
subgraph D1["Dimension 1: the 4 layers (where the signal occurs)"]
L1["Infrastructure: GPU, network, vector store, MCP server"]
L2["Model / inference: latency, tokens, saturation"]
L3["Application: RAG, prompt, orchestration, guardrails"]
L4["Agent / task: trajectory, tools, end-to-end success"]
L1 --> L2 --> L3 --> L4
end
subgraph D2["Dimension 2: the 4 axes (why you monitor)"]
A1["Reliability"]
A2["Availability"]
A3["Cost"]
A4["Compliance"]
end
subgraph D3["Dimension 3: the 5 questions (per signal)"]
Q["What -> Why -> How -> Threshold -> Impact if absent"]
end
subgraph D4["Dimension 4: maturity (the guide's levels 0 to 5)"]
M["0. Blind -> 1. Basic telemetry -> 2. Tracing -> 3. Online evaluation -> 4. Drift and feedback -> 5. Closed loop"]
end
D1 --> D2 --> D3 --> D4
Dimension 1: the four layers
Section titled “Dimension 1: the four layers”Every signal belongs to a layer. Structuring rule: the cascade. A low failure surfaces as a high symptom. Without correlation through a single trace_id across layers, you observe symptoms with no cause.
flowchart LR C2["Layer 2: provider rate-limit at 10:47"] --> C3["Layer 3: retries, prompt latency"] --> C4["Layer 4: task success rate dropping"] C4 -.->|"downward diagnosis via trace_id"| C2
The site describes the stack of a generative AI application in five layers in AI in cross-section: data, infrastructure, model, orchestration, application. The method’s four layers map onto them as follows (mind the word “application”, which does not mean the same thing):
| Method layer | AI in cross-section layer(s) |
|---|---|
| 1. Infrastructure: GPU, network, vector store, MCP server | infrastructure (GPUs, network, servers); the indexed content of the vector store belongs to data; the tools exposed by an MCP server are called from orchestration |
| 2. Model / inference: latency, tokens, saturation | model, at its boundary with infrastructure: TTFT, tokens per second and saturation are measured at the inference server, which the diagram places in infrastructure |
| 3. Application: RAG, prompt, orchestration, guardrails | orchestration (prompts, document retrieval); output guardrails also touch the application layer (policy violations, personal data in outputs) |
| 4. Agent / task: trajectory, tools, end-to-end success | orchestration for the trajectory and tool calls; application for the success the user perceives |
| no dedicated layer | data: source freshness and coverage, input distribution, which the method handles in the RAG grid of Part III |
Dimension 2: the four axes
Section titled “Dimension 2: the four axes”Reliability, availability, cost, compliance. Each component is read along these four axes in the grids of Part III.
Dimension 3: the five questions
Section titled “Dimension 3: the five questions”The template applied to every signal. You never instrument without answering “why” and “impact if absent”.
| Question | What it establishes |
|---|---|
| What | The precise signal: metric, trace, score, event |
| Why | The decision it informs. A signal with no decision is noise |
| How | The source and the instrument: OTel attribute, exporter, evaluator |
| Threshold | The SLI and the SLO, hence the alert trigger |
| Impact if absent | The failure that accumulates without this signal, and its severity |
Dimension 4: maturity
Section titled “Dimension 4: maturity”The method used to have its own five-level scale (Logging, Tracing, Evaluation, Monitoring, Closed loop). The site keeps only one, that of the guide, chapter 3, with six levels from 0 to 5. Correspondence:
| Guide level | What characterizes it | Former method level |
|---|---|---|
| 0. Blind operation | HTTP metrics only; cost discovered on the bill, quality through complaints | none: the method’s scale started at logging |
| 1. Basic telemetry | structured per-request logs (prompt, response, model, tokens, latency), aggregated cost, errors as metrics | 1. Logging |
| 2. Structured tracing | OpenTelemetry GenAI conventions, one span per logical step, queryable attributes | 2. Tracing |
| 3. Online evaluation | automated evaluations attached to traces and exposed as metrics, alertable quality regressions | 3. Evaluation |
| 4. Drift and feedback integration | embedding drift, batch RAG quality, user feedback correlated with traces | 4. Monitoring (approximate match: the method did not detail this level) |
| 5. Closed loop | failing traces become regression cases, changes gated on the regression set | 5. Closed loop |
The quality SLOs of Part IV assume at least level 3: without online evaluation there is no sample to measure.
In my view, many organizations stop at level 2 and believe they observe their system. They see what happened, not whether it was good. The value jump is at level 3, mastery at level 5.
3. The five signals
Section titled “3. The five signals”The signals are those of the guide, chapter 1: logs, metrics, traces, evaluations, user feedback, correlated by the trace ID. The method adds the typical storage of its reference configuration:
| Signal | Typical storage |
|---|---|
| Logs | VictoriaLogs, Loki |
| Metrics | VictoriaMetrics, Prometheus |
| Traces | OTLP backend (Tempo, Jaeger) |
| Evaluations | Langfuse, Arize Phoenix |
| User feedback | attached to the trace by its ID, next to the evaluation scores |
The “content events” (prompt, tool call, result) listed in the first version of the method are not a separate signal: they belong to logs and traces, as content capture (rule below).
4. The standard: OpenTelemetry GenAI
Section titled “4. The standard: OpenTelemetry GenAI”The instrumentation baseline to adopt: the OpenTelemetry GenAI semantic conventions, in Development status. Their versions, their history (including the move to a dedicated repository since 1.42), their attributes and content capture are detailed only once on the site, in the article Observing an LLM system, §2. The method relies mostly on gen_ai.request.model and gen_ai.response.model, gen_ai.provider.name, gen_ai.usage.input_tokens and gen_ai.usage.output_tokens (the basis of cost), gen_ai.response.finish_reasons (a first-order reliability signal) and gen_ai.operation.name. Its recording rules (Part II) start from the metrics of versions 1.40 and 1.41 (gen_ai.client.operation.duration, gen_ai.client.token.usage): see the revision note at the bottom of the page.
Content capture rule, aligned with the specification: instrumentations do not capture instructions, inputs and outputs by default; capturing them is an explicit opt-in. Three usage patterns: record nothing (the default); record the content on spans (gen_ai.system_instructions, gen_ai.input.messages, gen_ai.output.messages), reserved for cases where volume stays manageable and storage complies with regulation, for example in pre-production; store the content externally and record a reference on the span, the pattern the specification recommends in production. In a regulated environment the method retains the third, with masking at the collector. Content never goes into a metric label (see the guide, 9.6).
Revised on 2 October 2026: since version 1.42 (June 2026), the GenAI conventions live in the semantic-conventions-genai repository. At that date, its main branch, whose changes are not yet published in a release, lists gen_ai.client.operation.duration as recommended rather than required, and replaces gen_ai.client.token.usage with gen_ai.client.inference.usage.* counters (input, output, cache and reasoning tokens) and gen_ai.client.inference.operation.* histograms. Check which version your instrumentation emits before writing your rules.
Revised on 4 October 2026: signals aligned on the guide’s five (content events belong to logs and traces), correspondence table to the five layers of AI in cross-section, maturity scale replaced by the guide’s levels 0 to 5, OpenTelemetry conventions recap reduced to a summary with a pointer to the article, content capture rule aligned with the specification.