Skip to content

TechnicalExpert

Observing an LLM system: deployment playbooks for an open source stack (LLM, RAG, MCP)

For: engineers and SREs · architectsPrerequisites: Good knowledge of OpenTelemetry and the Collector, basic notions of LLMs, RAG and MCP.

Reading mode

This article is the “Go to production” step of the progression: it assumes the concepts of the guide and the practice of the labs; the GenAI method then takes over for quality SLOs, quantified impact and compliance. It is also the site’s reference for the OpenTelemetry GenAI semantic conventions, their history and the content capture rule (section 2), as well as for MCP instrumentation (section 7).

An application built on a language model rarely fails by returning a 500 code. It fails by returning a 200 OK, in 3.4 seconds, with a wrong answer, built from an outdated document, after a tool call that silently returned an empty list.

Classic APM sees the request. It sees none of the four decisions that produced the bad answer.

Question asked in operationsWhat classic APM saysWhat you need to instrument
Why is this answer wrong?nothingcontext content, retrieved documents, actual prompt
Why has the bill doubled?nothinginput and output tokens, per model and per route
Why has it been slow since Tuesday?overall HTTP latencybreakdown embeddings / search / reranking / inference
Did this tool really work?server 200 codeMCP call result, error.type, size of the response
Is quality drifting?nothingonline evaluation scores, correlated with the trace

Three distinct planes must be covered and a stack that covers only one leaves the other two in the dark:

  • the inference plane: model calls, tokens, latency, time to first token, cost;
  • the retrieval plane: vectorization, search, reranking, context assembly;
  • the tools plane: MCP calls, process boundary crossing, tool errors.

What follows is a series of deployment playbooks, from the minimal foundation to the production stack, with complete configurations. The building blocks are open source; for Langfuse, only the core is (see section 10).


2. The semantic foundation: OpenTelemetry GenAI

Section titled “2. The semantic foundation: OpenTelemetry GenAI”

The point that decides how long your stack will last is not the choice of backend, it is the vocabulary. This section is the site’s reference on these conventions: other content links here rather than retelling their history.

  • The gen_ai.* conventions have left the main repository: version 1.42.0 of the semantic conventions (12 June 2026) deprecated the gen_ai.*, openai.* and mcp.* attributes, metrics, events and spans and moved them to the open-telemetry/semantic-conventions-genai repository. The main repository carried on without them (1.43.0 in July 2026, 1.44.0 in August 2026).
  • That repository has no tagged release yet (checked on 4 October 2026). Its manifest declares a development schema (gen-ai-dev/1.42.0-dev) and builds on the general conventions 1.44.0. So you cannot pin a published version: you pin a commit.
  • No gen_ai.* attribute, span, metric or event is stable. All carry the Development status. The general attributes they reuse, such as error.type or network.transport, are stable.
  • The mcp.* attributes, introduced in 1.39.0, have followed the same path: deprecated in the main registry, they now live in the GenAI repository.

The renames already absorbed set the pace:

VersionChangeConsequence in operations
1.37.0 (August 2025)gen_ai.system → gen_ai.provider.name; per-message events (gen_ai.user.message, gen_ai.choice…) give way to the gen_ai.system_instructions, gen_ai.input.messages and gen_ai.output.messages attributes and the gen_ai.client.inference.operation.details eventdual reading mandatory during the transition; content masking rules to rewrite
1.38.0 (October 2025)gen_ai.evaluation.result eventquality scores enter the bus
1.39.0 (January 2026)MCP conventions (mcp.*)MCP tool calls become instrumentable as standard
1.40.0 (February 2026)retrieval spans and cache attributesRAG becomes instrumentable as standard
1.41.0 (April 2026)invoke_agent split into client and internal variants; tool name in the execute_tool span name; time-to-first-chunk metricsagent dashboards break
1.42.0 (June 2026)exit from the main repositoryno more published schema version to cite
GenAI repository, main branch, unreleasedgen_ai.client.token.usage and the gen_ai.token.type attribute replaced by gen_ai.client.inference.usage.* counters (broken down by gen_ai.token.modality) and two per-operation histograms; gen_ai.usage.cache_creation.input_tokens renamed gen_ai.usage.cache_write.input_tokensthe cost rules in section 5.4 read the old name: translate in the Collector the day your instrumentation changes

gen_ai.operation.name carries the nature of the operation. The values that matter to you in practice:

OperationSpan nameWhat it covers
chatchat {gen_ai.request.model}conversational inference call
embeddingsembeddings {gen_ai.request.model}vectorization of a query or a document
retrievalretrieval {gen_ai.data_source.id}search in a vector database or an index
execute_toolexecute_tool {gen_ai.tool.name}execution of a tool by the agent
invoke_agentinvoke_agent {gen_ai.agent.name}complete agent turn

The attributes to treat as the hard core:

INFERENCE
gen_ai.provider.name openai | anthropic | mistral_ai | aws.bedrock ...
(ollama, qdrant: custom values)
gen_ai.request.model requested model
gen_ai.response.model model actually served (may differ)
gen_ai.usage.input_tokens integer
gen_ai.usage.output_tokens integer
gen_ai.response.finish_reasons array: stop | length | tool_call | error ...
gen_ai.conversation.id conversation identifier
RETRIEVAL
gen_ai.data_source.id source identifier
gen_ai.retrieval.top_k number of documents requested
gen_ai.retrieval.query.text opt-in, sensitive content
gen_ai.retrieval.documents opt-in, sensitive content
MCP
mcp.method.name tools/call | resources/read | server/discover
mcp.protocol.version 2026-07-28 ...
mcp.session.id versions before 2026-07-28, HTTP only
network.transport pipe (stdio) | tcp (HTTP)
error.type set only on failure

Four clarifications, checked against the conventions repository on 4 October 2026:

  • gen_ai.provider.name has a list of well-known values (openai, anthropic, mistral_ai, aws.bedrock, azure.ai.openai, gcp.vertex_ai, cohere, deepseek, groq…): if one applies, it must be used; otherwise a custom value is allowed. ollama is not one of them: it is a legitimate custom value, like qdrant on the retrieval span in playbook 3. Watch out for Mistral AI, whose value is mistral_ai, not mistral.
  • gen_ai.response.finish_reasons is an array, one reason per returned choice, with no closed list in the registry (examples: stop, length, error). The output messages schema lists stop, length, content_filter, tool_call, compaction and error: it is indeed tool_call, singular, not the raw tool_calls value some providers return. A reason that was expected but not received (failed generation, interrupted stream) must be reported as error.
  • mcp.method.name follows the same well-known values rule. The registry still reflects the methods that predate the MCP 2026-07-28 specification (initialize is among them): server/discover is a custom value there for now.
  • Metrics: two cover most of the day-to-day need, in my view. gen_ai.client.operation.duration (histogram, in seconds) gives latency per operation. For tokens, until they moved, the conventions described the gen_ai.client.token.usage histogram, broken down by gen_ai.token.type; the main branch of the GenAI repository replaces it with the gen_ai.client.inference.usage.input_tokens and gen_ai.client.inference.usage.output_tokens counters. Check which one your instrumentation emits: the rules in section 5.4 are written on gen_ai.client.token.usage, which is exactly the case where translation in the Collector protects you. For time to first chunk when streaming, gen_ai.client.operation.time_to_first_chunk has existed since 1.41.0.

Enabling it on the SDK side:

Fenêtre de terminal
export OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental

Not all SDKs honor this flag in the same way. Check what actually comes out before building dashboards on it: the acceptance command is given in playbook 1.

Prompts, system instructions and responses are the most sensitive data in the stack. The conventions are explicit: instrumentations should not capture this content by default and should offer an explicit opt-in. They describe three modes:

  1. by default, no capture;
  2. on spans, in the gen_ai.system_instructions, gen_ai.input.messages and gen_ai.output.messages attributes, all at the Opt-In level: suited when volume stays manageable and either privacy regulations do not apply or the telemetry storage complies with them, for example in pre-production;
  3. in external storage, with only a reference on the span: the mode recommended in production, because it allows separate access controls.

The same content can also travel in the gen_ai.client.inference.operation.details event, also Opt-In, which stores it independently from traces. The gen_ai.retrieval.query.text, gen_ai.retrieval.documents and gen_ai.tool.definitions attributes are Opt-In too.

The rule adopted across the whole site, whose wording other content reuses: content capture is disabled by default in production; it is enabled explicitly, per environment, with its own retention; preferably outside the indexed span attributes (dedicated event or log, or referenced external storage), with masking at the Collector; and its absence is verified by an actual search in each backend. The implementation, application barrier and Collector barrier, is detailed in section 9.

The gen_ai.* namespace is reserved for the conventions. Attributes specific to an application use a namespace of its own: in these playbooks llm.* and rag.*, elsewhere on the site mttl.* or app.*, never gen_ai.*. A home-grown attribute placed under gen_ai.* would collide with the next version of the convention.


One bus, several consumers. This is, in my view, the most structuring architecture decision.

+---------------------------------------------------------------+
| APPLICATION: agent, RAG chain, MCP client |
| OpenTelemetry SDK + GenAI instrumentation |
+---------------------------------------------------------------+
v OTLP 4317 / 4318
+---------------------------------------------------------------+
| OTEL COLLECTOR: gateway |
| normalization > redaction > sampling > derivation |
+---------------------------------------------------------------+
v v v v
+---------------------------------------------------------------+
| Tempo/Jaeger Prometheus/VM Loki/OpenSearch Langfuse |
| traces metrics logs LLM + evals |
+---------------------------------------------------------------+
v
+---------------------------------------------------------------+
| GRAFANA: operations, cost, quality |
+---------------------------------------------------------------+

The non-negotiable principle: the application emits only OTLP. It does not know that Langfuse, Tempo or Prometheus exist. The day you change LLM backend, you do not redeploy forty services: you change the exporters section of the Collector.

The symmetrical mistake, to be avoided, is to run two instrumentation SDKs side by side, the LLM observability vendor’s and OpenTelemetry’s. You get two disjoint trace trees, double ingestion billing and trace identifiers that do not correlate. One SDK, one gateway.


4. Playbook 1: Minimal foundation in half a day

Section titled “4. Playbook 1: Minimal foundation in half a day”

Goal: see the first complete LLM trace, with tokens and cost, without deciding anything final.

RoleBuilding blockLicense
InstrumentationOpenLIT or OpenLLMetry (Traceloop)Apache-2.0
GatewayOpenTelemetry Collector contribApache-2.0
LLM backendLangfuse (self-hosted)MIT (core)

Langfuse v3 is not a single binary: it needs PostgreSQL, ClickHouse, Redis and S3-compatible object storage. Do not rebuild this composition by hand, start from the maintained one:

Fenêtre de terminal
git clone --depth 1 https://github.com/langfuse/langfuse.git
Fenêtre de terminal
cd langfuse && docker compose up -d

The interface is on http://localhost:3000. Create a project, retrieve the pk-lf-... / sk-lf-... key pair, then build the OTLP authentication header:

Fenêtre de terminal
export LANGFUSE_BASIC_AUTH=$(printf '%s:%s' "$LANGFUSE_PUBLIC_KEY" "$LANGFUSE_SECRET_KEY" | base64 -w0)

Langfuse exposes a native OTLP endpoint on /api/public/otel, over HTTP only, no gRPC. The Collector absorbs this constraint: the application can keep speaking gRPC if it wishes.

otel-collector.yaml:

receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
memory_limiter:
check_interval: 1s
limit_percentage: 80
spike_limit_percentage: 20
batch:
timeout: 2s
send_batch_size: 512
# Dual reading: gen_ai.system (< 1.37) -> gen_ai.provider.name
transform/semconv:
error_mode: ignore
trace_statements:
- set(span.attributes["gen_ai.provider.name"], span.attributes["gen_ai.system"])
where span.attributes["gen_ai.provider.name"] == nil
and span.attributes["gen_ai.system"] != nil
exporters:
otlphttp/langfuse:
endpoint: http://langfuse-web:3000/api/public/otel
headers:
Authorization: "Basic ${env:LANGFUSE_BASIC_AUTH}"
debug:
verbosity: detailed
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, transform/semconv, batch]
exporters: [otlphttp/langfuse, debug]
Fenêtre de terminal
pip install openlit
import openlit
openlit.init(
application_name="assistant-support",
environment="lab",
otlp_endpoint="http://otel-collector:4318",
capture_message_content=False, # read section 9 before switching to True
)

Environment variables, to be set at deployment level rather than in the code:

Fenêtre de terminal
OTEL_SERVICE_NAME=assistant-support
OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental
OTEL_RESOURCE_ATTRIBUTES=deployment.environment.name=lab,service.version=1.4.2

Do not validate on “I can see something in the interface”. Validate on the attributes:

Fenêtre de terminal
docker compose logs otel-collector | grep -E 'gen_ai\.(provider\.name|request\.model|usage\.)' | head -20
CheckExpected
One span per model callname chat <model>
Tokens presentgen_ai.usage.input_tokens and output_tokens non-zero
Normalized providergen_ai.provider.name set, including on older SDKs
Single traceall spans of a request share the same trace_id
No contentneither prompt nor completion in the debug output

If the last point fails, stop there and deal with section 9 before going any further.


The playbook 1 foundation does not hold up in production: it sends 100% of traces, derives no metrics and has only one consumer.

Tail sampling requires that all spans of a trace land on the same Collector instance. With several replicas, this calls for a first tier that distributes by trace_id.

Whichever agent receives a span, the routing key sends it to the gateway that holds its trace:

flowchart TB
  APP["Applications<br/>OTLP 4317 / 4318"]
  subgraph E1["Tier 1: agents"]
    A1["Agent 1<br/>memory_limiter, batch,<br/>load_balancing"]
    A2["Agent 2<br/>memory_limiter, batch,<br/>load_balancing"]
  end
  subgraph E2["Tier 2: gateways"]
    G1["Gateway A<br/>tail_sampling, redaction,<br/>connectors"]
    G2["Gateway B<br/>tail_sampling, redaction,<br/>connectors"]
  end
  APP --> A1
  APP --> A2
  A1 -->|"trace X"| G1
  A2 -->|"trace X"| G1
  A1 -->|"trace Y"| G2
  A2 -->|"trace Y"| G2
  G1 --> TP["Tempo<br/>traces"]
  G1 --> VM["VictoriaMetrics<br/>metrics"]
  G1 --> LF["Langfuse<br/>LLM"]
  G2 --> TP
  G2 --> VM
  G2 --> LF

Tier 1, the essentials:

exporters:
# formerly "loadbalancing", name kept as a deprecated alias
load_balancing:
routing_key: traceID
protocol:
otlp:
tls:
insecure: true
resolver:
dns:
hostname: otel-gateway-headless.observability.svc.cluster.local
port: 4317

In my view, random sampling is a bad choice here: it throws away as many incidents as nominal traffic. The decision must be made once the trace is complete.

processors:
tail_sampling:
decision_wait: 30s
num_traces: 100000
expected_new_traces_per_sec: 500
policies:
- name: erreurs
type: status_code
status_code:
status_codes: [ERROR]
- name: outil-en-echec
type: string_attribute
string_attribute:
key: error.type
values: [".+"]
enabled_regex_matching: true
- name: latence
type: latency
latency:
threshold_ms: 8000
- name: generation-anormalement-longue
type: numeric_attribute
numeric_attribute:
key: gen_ai.usage.output_tokens
min_value: 2000
# min_value: 0 and max_value: 0 would be ignored (0 means "not set"
# for numeric_attribute): we use an OTTL condition instead.
- name: recuperation-vide
type: ottl_condition
ottl_condition:
error_mode: ignore
span:
- span.attributes["rag.documents.count"] == 0
- name: nominal
type: probabilistic
probabilistic:
sampling_percentage: 5

The policies are combined with a logical OR: a trace selected by any of them is kept. The intended result: 100% of the pathological, 5% of the rest.

Rather than instrumenting twice, let the Collector build the RED metrics from the spans:

processors:
# gen_ai.response.finish_reasons is an array (one reason per choice):
# we extract a string from it, into an attribute of the internal model.
transform/finish_reason:
error_mode: ignore
trace_statements:
- set(span.attributes["llm.finish_reason"], span.attributes["gen_ai.response.finish_reasons"][0])
where IsList(span.attributes["gen_ai.response.finish_reasons"])
- set(span.attributes["llm.finish_reason"], span.attributes["gen_ai.response.finish_reasons"])
where IsString(span.attributes["gen_ai.response.finish_reasons"])
connectors:
spanmetrics:
histogram:
explicit:
buckets: [100ms, 250ms, 500ms, 1s, 2s, 5s, 10s, 30s, 60s]
dimensions:
- name: gen_ai.provider.name
- name: gen_ai.request.model
- name: gen_ai.operation.name
- name: mcp.method.name
- name: deployment.environment.name
- name: llm.finish_reason

The connector copies the attribute value as is: an array would stay an array and the label obtained after translation to Prometheus would be a serialized form (such as ["length"]) that the ="length" filter would not find. Hence the transform/finish_reason processor, which extracts a string from it. It must appear in the traces pipeline before the spanmetrics exporter. It keeps only the reason of the first choice, which is enough for calls with a single completion (n=1), the common case. The attribute produced belongs to the internal model (llm.*) and not to gen_ai.*, following the rule in section 2.4.

Watch the cardinality. Each dimension multiplies the number of active series. The six above are bounded. gen_ai.conversation.id, a user identifier or the name of a retrieved document are not: these values belong in traces, never in metric labels. The “Cost is decided at the Collector” simulator puts figures on the effect of these choices.

Prometheus does not know your rates. Expose them as a static series, through the textfile collector of node_exporter or a twenty-line exporter. Prices change fast and vary by contract: this article does not give any; use your provider’s dated price list, converted into your currency at a date you record. Replace each <...> with a numeric value in euros per token:

genai_prix.prom
# To be filled in from your provider's dated price list (euros per token)
genai_prix_euro_par_token{model="grand-modele",token_type="input"} <prix_entree_grand>
genai_prix_euro_par_token{model="grand-modele",token_type="output"} <prix_sortie_grand>
genai_prix_euro_par_token{model="petit-modele",token_type="input"} <prix_entree_petit>
genai_prix_euro_par_token{model="petit-modele",token_type="output"} <prix_sortie_petit>

As long as this file does not exist, the join below returns nothing: genai:cout_euro_par_heure is not produced, but genai:jetons:rate5m is and you can alert on tokens with the same rule comparing to the seven-day baseline.

Then the recording rules:

groups:
- name: genai-cout
interval: 60s
rules:
- record: genai:jetons:rate5m
expr: |
sum by (service_name, gen_ai_request_model, gen_ai_token_type) (
rate(gen_ai_client_token_usage_sum[5m])
)
- record: genai:cout_euro_par_heure
expr: |
sum by (service_name, gen_ai_request_model) (
genai:jetons:rate5m * 3600
* on (gen_ai_request_model, gen_ai_token_type) group_left()
label_replace(
label_replace(genai_prix_euro_par_token,
"gen_ai_request_model", "$1", "model", "(.*)"),
"gen_ai_token_type", "$1", "token_type", "(.*)")
)

The alerts that are actually useful:

- alert: CoutHoraireLLMAnormal
expr: |
genai:cout_euro_par_heure
> 3 * avg_over_time(genai:cout_euro_par_heure[7d] offset 1h)
for: 15m
annotations:
summary: "Cout LLM x3 vs reference 7 jours"
- alert: TroncatureDeGeneration
expr: |
sum(rate(traces_span_metrics_calls_total{
llm_finish_reason="length"}[10m]))
/ sum(rate(traces_span_metrics_calls_total{
gen_ai_operation_name="chat"}[10m])) > 0.05
for: 10m
annotations:
summary: "Plus de 5% des generations coupees par la limite de jetons"

The second alert relies on the llm.finish_reason dimension added in 5.3. In my view, it is rarely set up even though it can explain part of the “incomplete answers” reported by users.


FailureUser symptomSignal that reveals itSignal that does not reveal it
Empty or off-topic retrieval“it makes things up”rag.documents.count, top-1 scorelatency, HTTP code
Context retrieved but ignored“it misses the point”ratio of cited / retrieved documentseverything else
Stale index“it gives the old procedure”rag.index.version, index agemodel quality
Truncated contextpartial answerrag.context.truncated, finish_reasons=lengthretrieval scores
Embedder driftslow, diffuse degradationscore distribution over timean isolated trace

The last two rows are invisible on a single trace. They only appear in the distribution, so in metrics, not in traces.

The chain splits into three spans, not one.

from opentelemetry import trace
tracer = trace.get_tracer("rag.pipeline")
INDEX = "kb-support-fr"
INDEX_VERSION = "v7"
MODELE = "grand-modele" # illustrative name
def repondre(question: str) -> str:
with tracer.start_as_current_span("embeddings text-embedding-3-large") as sp:
sp.set_attribute("gen_ai.operation.name", "embeddings")
sp.set_attribute("gen_ai.provider.name", "openai")
sp.set_attribute("gen_ai.request.model", "text-embedding-3-large")
vecteur = embedder.encode(question)
with tracer.start_as_current_span(f"retrieval {INDEX}") as sp:
sp.set_attribute("gen_ai.operation.name", "retrieval")
sp.set_attribute("gen_ai.data_source.id", INDEX)
sp.set_attribute("gen_ai.provider.name", "qdrant")
sp.set_attribute("gen_ai.retrieval.top_k", 8)
sp.set_attribute("rag.index.version", INDEX_VERSION)
docs = store.search(vecteur, k=8)
sp.set_attribute("rag.documents.count", len(docs))
sp.set_attribute("rag.score.top", docs[0].score if docs else 0.0)
sp.set_attribute("rag.score.min_kept", docs[-1].score if docs else 0.0)
# identifiers, never the content
sp.set_attribute("rag.document.ids", [d.id for d in docs][:8])
with tracer.start_as_current_span(f"chat {MODELE}") as sp:
contexte, tronque = assembler(docs, budget=12_000)
sp.set_attribute("rag.context.chars", len(contexte))
sp.set_attribute("rag.context.truncated", tronque)
reponse = client.chat(question, contexte)
sp.set_attribute("rag.documents.cited", compter_citations(reponse, docs))
return reponse

Three choices deserve an explanation.

rag.index.version on every retrieval span. Without this attribute, a failed reindexing cannot be detected after the fact: you will see a drop in quality without being able to date it or attribute it. With it, a single query is enough to confirm the correlation.

Document identifiers, not their content. An identifier is bounded, not personal and enough to replay the retrieval. The content of the chunk is bulky, often confidential and has no place in a span by default. The gen_ai.retrieval.documents attribute exists in the convention, where it is classified as opt-in; this classification is a warning, not a formality.

rag.documents.cited. In my view, this is the most useful measure in the whole chain. The cited / count ratio measures actual use of the context. A ratio that collapses while retrieval scores stay good signals that the problem has moved from the retriever to the generator, typically after a change of model or system prompt.

Span attributes do not become metrics on their own. The rag_* series below are application metrics that you emit yourself, with the OpenTelemetry SDK, alongside the span attributes:

from opentelemetry import metrics
meter = metrics.get_meter("rag.pipeline")
documents = meter.create_histogram(
"rag.documents.count",
explicit_bucket_boundaries_advisory=[0, 1, 2, 4, 8])
score_top = meter.create_histogram(
"rag.score.top",
explicit_bucket_boundaries_advisory=[0.2, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0])
assemblages = meter.create_counter("rag.context.truncated")
# in repondre(), after the search and then after context assembly
etiquettes = {"rag.index.version": INDEX_VERSION}
documents.record(len(docs), etiquettes)
score_top.record(docs[0].score if docs else 0.0, etiquettes)
assemblages.add(1, {**etiquettes, "truncated": str(tronque).lower()})

Once exported over OTLP to a Prometheus-compatible backend, with the usual name translation (dots replaced by underscores, _total suffix for counters, no declared unit), they give the following queries:

sum(rate(rag_documents_count_bucket{le=~"0(\\.0)?"}[5m]))
/ sum(rate(rag_documents_count_count[5m])) -> empty retrieval rate
histogram_quantile(0.5, sum by (le, rag_index_version) (
rate(rag_score_top_bucket[30m]))) -> median top-1 score
sum(rate(rag_context_truncated_total{truncated="true"}[15m]))
/ sum(rate(rag_context_truncated_total[15m])) -> truncation rate

A drop in the median top-1 score that coincides with a change in rag_index_version is a diagnosis, not a hypothesis.


An MCP server over stdio is a separate process that communicates in JSON-RPC over pipes. There are no HTTP headers, so no traceparent where OpenTelemetry propagators usually look for it. Without specific handling, every tool call produces an orphan trace: you have the latency, you have lost the link with the conversation that triggered it.

The answer is standardized. SEP-414 of the MCP protocol, integrated into the 2026-07-28 specification, documents the use of params._meta as the carrier of the W3C context, with an explicit exception to the DNS prefixing rule for keys.

{
"jsonrpc": "2.0",
"id": 2,
"method": "tools/call",
"params": {
"name": "get_weather",
"arguments": { "location": "New York" },
"_meta": {
"traceparent": "00-0af7651916cd43dd8448eb211c80319c-00f067aa0ba902b7-01"
}
}
}

The keys are traceparent, tracestate and baggage, without a prefix. An implementation that wrote io.modelcontextprotocol.traceparent would break correlation with the whole existing ecosystem.

MCP CLIENT: agent MCP SERVER: tool
+-------------------------------+ +-------------------------------+
| span kind = CLIENT | | span kind = SERVER |
| name: tools/call get_weather | | name: tools/call get_weather |
| mcp.method.name = tools/call | | mcp.method.name = tools/call |
| mcp.protocol.version | | mcp.protocol.version |
| = 2026-07-28 | | = 2026-07-28 |
| network.transport = pipe | | network.transport = pipe |
+-------------------------------+ +-------------------------------+
JSON-RPC: params._meta.traceparent = 00-<trace>-<span>-01
same trace_id on both sides, parent-child relationship preserved

The same mechanism, unfolded over time across a complete agent turn:

sequenceDiagram
  participant U as User
  participant A as Agent, MCP client
  participant L as LLM
  participant S as MCP server
  U->>A: question
  A->>L: chat, CLIENT span
  L-->>A: finish_reason tool_call
  A->>S: tools/call with params._meta.traceparent
  Note over A,S: same trace_id, the SERVER span is a child of the CLIENT span
  S-->>A: result, or error with error.type
  A->>L: chat with the tool result
  L-->>A: final answer
  A-->>U: answer
from opentelemetry import trace
from opentelemetry.propagate import inject
from opentelemetry.trace import SpanKind
tracer = trace.get_tracer("mcp")
VERSION_MCP = "2026-07-28"
def appeler_outil(message: dict, serveur: str):
methode = message["method"]
with tracer.start_as_current_span(
f"{methode} {serveur}", kind=SpanKind.CLIENT
) as sp:
sp.set_attribute("mcp.method.name", methode)
sp.set_attribute("mcp.protocol.version", VERSION_MCP)
sp.set_attribute("network.transport", "pipe")
meta = message.setdefault("params", {}).setdefault("_meta", {})
# since 2026-07-28, every request carries its protocol version
# and the client capabilities (two mandatory keys)
meta["io.modelcontextprotocol/protocolVersion"] = VERSION_MCP
meta.setdefault("io.modelcontextprotocol/clientCapabilities", {})
inject(meta) # traceparent WITHOUT DNS prefix, SEP-414
reponse = transport.send(message)
if "error" in reponse:
sp.set_attribute("error.type", str(reponse["error"].get("code")))
sp.set_status(trace.StatusCode.ERROR)
return reponse
from opentelemetry import trace
from opentelemetry.propagate import extract
from opentelemetry.trace import SpanKind
tracer = trace.get_tracer("mcp")
def traiter(message: dict):
porteur = (message.get("params") or {}).get("_meta") or {}
ctx = extract(porteur) # reads traceparent / tracestate / baggage
methode = message["method"]
with tracer.start_as_current_span(
f"{methode} {NOM_SERVEUR}", context=ctx, kind=SpanKind.SERVER
) as sp:
sp.set_attribute("mcp.method.name", methode)
version = porteur.get("io.modelcontextprotocol/protocolVersion")
if version:
sp.set_attribute("mcp.protocol.version", version)
sp.set_attribute("network.transport", "pipe")
if methode == "tools/call":
sp.set_attribute("gen_ai.tool.name", message["params"]["name"])
return executer(message)

If your client or server uses an official MCP SDK, part of this work may already be done upstream and third-party instrumentations also exist. Check before instrumenting by hand: double instrumentation produces duplicate spans and skews every count.

Sessions and handshake. The MCP 2026-07-28 specification makes the protocol stateless: it removes the initialize handshake, protocol sessions and the Mcp-Session-Id header. Every request carries its protocol version and the client capabilities in _meta and a server announces the versions it accepts through the server/discover method. The session identifier never appeared in the JSON-RPC message anyway: in earlier versions, it was an HTTP header of the Streamable HTTP transport, absent over stdio. If you still have to serve clients older than 2026-07-28 over HTTP, read mcp.session.id from that header, not from the message.

IndicatorBreakdownWhy
Failure rate per toolerror.type by gen_ai.tool.namea tool that fails 30% of the time degrades the agent with no visible error
p95 latency per toolmcp.method.name="tools/call"a slow tool is multiplied by the number of agent turns
Calls per agent turnMCP spans under invoke_agentdetection of tool call loops
Tools called outside the listgen_ai.tool.name outside the allowlistprompt drift or indirect injection
Requests rejected for protocol versionUnsupportedProtocolVersion errors by mcp.protocol.versionclients or servers not yet migrated to 2026-07-28

The fourth row deserves a word. An MCP server is an execution surface reachable by the content the model processes. MCP observability is therefore not only a performance topic: the list of tools actually called, compared with the list of tools expected for a given route, is, in my view, an important security signal. It is also one of the few places where an indirect injection can become visible.


Without a quality measure, the stack tells you everything is fine while relevance collapses.

The two stages of the loop, offline before deployment and online after, answer each other:

flowchart TB
  CH["Change<br/>prompt, model,<br/>chunking, index"] --> CI{"Offline in CI<br/>reference set"}
  CI -->|"below threshold"| BL["Deployment<br/>blocked"]
  CI -->|"thresholds met"| PR["Production"]
  PR --> EC["1 to 3% sample<br/>pinned judge"]
  EC --> EV["gen_ai.evaluation.result<br/>correlated with the span"]
  EV --> AL{"Drift<br/>alert"}
  AL -->|"yes"| DG["Diagnosis<br/>index, prompt, model"]
  DG --> CH
  DG -.->|"new cases"| CI

A versioned reference set, run on every change of prompt, model, chunking or index.

Fenêtre de terminal
pip install "ragas==0.4.3"
import math
import statistics
from ragas import evaluate
# Legacy import, accepted by evaluate() in 0.4.3 with a deprecation
# warning (the new metrics live in ragas.metrics.collections).
from ragas.metrics import (
faithfulness, answer_relevancy,
context_precision, context_recall,
)
resultat = evaluate(
dataset=jeu_de_reference, # question, context, answer, ground truth
metrics=[faithfulness, answer_relevancy,
context_precision, context_recall],
)
SEUILS = {"faithfulness": 0.85, "context_recall": 0.80}
for metrique, seuil in SEUILS.items():
# resultat[metrique] is the list of per-example scores: we compare the mean
scores = [s for s in resultat[metrique] if s is not None and not math.isnan(s)]
moyenne = statistics.fmean(scores) if scores else 0.0
if moyenne < seuil:
raise SystemExit(f"REGRESSION {metrique}: {moyenne:.3f} < {seuil}")

The 0.85 and 0.80 thresholds are examples, to be calibrated on your own measurements. The four metrics do not measure the same thing and the diagnosis depends on which one drops:

Metric droppingComponent at fault
context_recallchunking, embedder, top_k too low
context_precisionno reranking, score threshold too permissive
faithfulnesssystem prompt, model, truncated context
answer_relevancysystem prompt, rephrasing of the question

The reference set does not contain the questions your users will ask tomorrow. You need to judge a sample of real traffic. The site’s common principle, detailed in the guide, section 5.7: what counts is the number of judgments per segment and per period, not the percentage. A rate of 1 to 3% is an example starting point, to be adjusted to get enough judgments per alert window; since each judgment is itself a model call, cost quickly limits this rate.

The result is written to the bus as a gen_ai.evaluation.result event, introduced in 1.38.0, correlated with the original span. An event does not become a metric on its own: the evaluation service also emits an application histogram, for example meter.create_histogram("genai.evaluation.score") recorded with the evaluation_name attribute, whose bucket boundaries include the quality threshold (for example explicit_bucket_boundaries_advisory=[0.5, 0.6, 0.7, 0.8, 0.9, 1.0]). It thus becomes an alertable metric.

The alert is on a proportion, not an average: the share of judged responses whose score is above the threshold, compared with a target. An average can stay flat while a growing fraction of responses falls below the threshold.

- alert: SLOQualiteFidelite
expr: |
(
1 - sum(rate(genai_evaluation_score_bucket{evaluation_name="faithfulness", le="0.8"}[6h]))
/ sum(rate(genai_evaluation_score_count{evaluation_name="faithfulness"}[6h]))
) < 0.95
and sum(increase(genai_evaluation_score_count{evaluation_name="faithfulness"}[6h])) >= 100
for: 30m
annotations:
summary: "Under 95% of judged responses above 0.8 faithfulness over 6h"

The threshold (0.8), the target (95%) and the minimum of 100 judgments are examples. This rule still compares a point value with the target: the GenAI method, part IV compares the lower bound of the confidence interval (Wilson bound), sizes the sample and derives error budget and multi-window burn rate from it. It is the reference for quality SLOs.

Three precautions, in order of importance:

  1. Pin the judge model version. A judge that changes version shifts all your scores and you will spend a week looking for an application regression that does not exist.
  2. Do not judge with the model being judged. A judge model tends to favor its own outputs: this self-preference bias was measured by Panickssery, Bowman and Feng (NeurIPS 2024) on summarization tasks (XSUM, CNN/DailyMail), with a magnitude that varies across models.
  3. Budget for the judge. For example, at 3% of 500,000 requests per day, that makes 15,000 daily calls, which will show up in the playbook 2 cost dashboard. Label them with a distinct service.name so as not to confuse them with application traffic.

9. Confidentiality: do not turn observability into a leak

Section titled “9. Confidentiality: do not turn observability into a leak”

This is the point where an LLM observability stack becomes an incident.

Prompts and completions contain, by construction, what the user wrote: personal data, health data, contractual elements, sometimes a secret pasted into a chat window. The documents retrieved by RAG contain the internal knowledge base, including what the user had no right to see. All of this is replicated, indexed and retained.

DataRiskWhere to handle it
User promptpersonal data, secretscapture disabled by default, masking at the Collector
Completionrestitution of context datasame treatment, short retention
Retrieved documentsknowledge base leak, bypassing access rightsidentifiers only, never the content
MCP call argumentsidentifiers, tokens, pathsargument allowlist, no bulk capture
gen_ai.conversation.idre-identification by cross-referencingpseudonymize, never as a metric label

The rule, set out in section 2.3: content is disabled by default and enabled explicitly, per environment, with its own retention. The flag exists on the SDK side (capture_message_content, or an OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT variable depending on the implementation; the name still varies, check yours). It is not enough: a flag gets set back to true by accident, in the middle of a debugging session. The second barrier is in the Collector and that one cannot be bypassed from an application.

processors:
transform/confidentialite:
error_mode: ignore
trace_statements:
# 1. Masking of the most common patterns
- replace_pattern(span.attributes["gen_ai.input.messages"],
"[\\w.+-]+@[\\w-]+\\.[\\w.]{2,}", "[courriel]")
- replace_pattern(span.attributes["gen_ai.input.messages"],
"\\b(?:\\d[ -]?){13,19}\\b", "[carte]")
- replace_pattern(span.attributes["gen_ai.output.messages"],
"[\\w.+-]+@[\\w-]+\\.[\\w.]{2,}", "[courriel]")
# 2. Deletion outside authorized environments
- delete_key(span.attributes, "gen_ai.input.messages")
where resource.attributes["deployment.environment.name"] == "production"
- delete_key(span.attributes, "gen_ai.output.messages")
where resource.attributes["deployment.environment.name"] == "production"
- delete_key(span.attributes, "gen_ai.retrieval.documents")

Depending on the version of your instrumentation, the content may arrive not as a span attribute but as a log record, the GenAI events. replace_pattern only acts on a string value: if your instrumentation emits messages in structured form, masking does not apply and only deletion protects you. Finally, duplicate the same rules in log_statements, replacing the span. prefix with log.: a clean traces pipeline paired with an unfiltered logs pipeline is, in my view, the easiest trap to miss in this section.


10. Choosing the building blocks: the licensing question

Section titled “10. Choosing the building blocks: the licensing question”

The tool landscape and its licenses have a single home on the site: chapter 6 of the guide. Only the points the playbooks depend on remain here.

Playbook building blockLicensePoint of attention
OpenTelemetry Collector, Jaeger, Prometheus, OpenLIT, OpenLLMetry, RagasApache-2.0none for internal use
VictoriaMetricsApache-2.0 (community edition, cluster version included)downsampling, multiple retentions, anomaly detection in the enterprise edition
Grafana, Tempo, LokiAGPL-3.0modifying the component and opening it to users over the network, even without distributing it, obliges you to offer them the modified source code
Langfuse (playbook 1)MIT for the coresome features under a commercial license; owned by ClickHouse since January 2026 (ClickHouse announcement)
Arize Phoenix (kit of the labs)Elastic License 2.0 (source available)not open source in the OSI sense: prohibited from being offered as a managed service to third parties; no effect for internal use

To choose the building block for the LLM plane, apply the same three criteria to every option: the license, the ability to self-host and the ability to receive OTLP with the gen_ai.* conventions without a proprietary SDK. In every case, keep your existing stack for the rest: you do not duplicate Tempo and Prometheus, you add a consumer to the bus.


The criterion I recommend keeping above all others: start from a user complaint and trace it back to the cause in less than five minutes, without access to the production machine.

  • A user request produces a single trace, from the HTTP entry point up to and including the MCP server.
  • Input and output tokens are present on 100% of inference spans.
  • Hourly cost per service and per model is displayed and alerted on deviation from the weekly baseline.
  • Retrieval spans carry rag.documents.count, the top-1 score and the index version.
  • The ratio of cited documents to retrieved documents is measured.
  • MCP client and server spans share the same trace_id via params._meta.traceparent.
  • The failure rate is visible per MCP tool, not only in aggregate.
  • The rate of generations cut off by the token limit is tracked.
  • Prompts and completions are absent from production backends, verified by an actual search.
  • An online quality score exists, fed by a sample, with a pinned judge model.
  • Sampling keeps 100% of errors and slow traces.
  • No unbounded attribute is used as a metric label.
  • Switching from one backend to another requires no application change.

If the last box is not ticked, you have not deployed an observability stack: you have deployed a proprietary client with extra steps.


Three ideas to remember, in this order.

Vocabulary comes before tooling. The GenAI conventions are on the move: repository relocated, no stable attribute, regular renames. Instrument towards OpenTelemetry, but build your dashboards and alerts on an internal model that you control and absorb the changes in the Collector.

RAG and MCP are the two planes that stacks forget. Inference is instrumented by default by every library. Retrieval and tool calls are not and they are, in my view, what produces a good share of the bad answers. A retrieval score, an index version and a traceparent in _meta: three elements, half a day of work as an order of magnitude and a large part of incidents becomes diagnosable.

Observing an LLM system is processing personal data. A complete trace contains what the user wrote and what the internal knowledge base contained. Content capture is therefore decided as a compliance measure, with one barrier on the application side and a second on the Collector side, tested by an actual search in the backends. A stack that exposes prompts to the whole operations team is not progress in observability, it is an incident waiting to be declared.

Next in the progression: VictoriaMetrics for LLMs for the production metrics backend, then the GenAI method for quality SLOs, quantified impact and compliance. To go back to the concepts: the guide; to practice on a complete stack: the labs.

Revised on 2 October 2026: prices removed, MCP section updated for the 2026-07-28 specification (no more protocol sessions or initialize handshake, protocol version in _meta), llm.finish_reason dimension added to the spanmetrics connector, rag_* and genai_evaluation_score metrics presented as application metrics to emit, Ragas gate corrected and pinned at 0.4.3, exporter renamed load_balancing, “empty retrieval” policy switched to an OTTL condition, OTTL statements written with prefixed paths, prices made illustrative, VictoriaMetrics license corrected, acquisition of Langfuse by ClickHouse mentioned, choice of the LLM building block reworded around explicit criteria with alternatives, scope of the AGPL clarified, judges’ self-preference bias sourced, estimates requalified as opinions or orders of magnitude, clientCapabilities key added to the MCP client, le filter adapted to Prometheus 3 normalization.

Revised on 4 October 2026: page moved to MDX with the progression box, section 2 made the site’s reference on the OpenTelemetry GenAI conventions (version history dated and checked against the conventions repository, well-known values of gen_ai.provider.name with mistral_ai and ollama presented as a custom value, tool_call singular in gen_ai.response.finish_reasons, ongoing replacement of gen_ai.client.token.usage, execute_tool and invoke_agent span names, content capture rule, namespace for home-grown attributes), quality alert reworded as a proportion with a link to the method, judge sampling tied to the guide’s principle, licensing section reduced to the playbooks’ building blocks with a link to the guide, links to the guide, the labs and the method.