Technical
Observing an LLM system: deployment playbooks for an open source stack (LLM, RAG, MCP)
For: engineers and SREs · architectsPrerequisites: Good knowledge of OpenTelemetry and the Collector, basic notions of LLMs, RAG and MCP.
This article is the “Go to production” step of the progression: it assumes the concepts of the guide and the practice of the labs; the GenAI method then takes over for quality SLOs, quantified impact and compliance. It is also the site’s reference for the OpenTelemetry GenAI semantic conventions, their history and the content capture rule (section 2), as well as for MCP instrumentation (section 7).
1. Framing the problem correctly
Section titled “1. Framing the problem correctly”An application built on a language model rarely fails by returning a 500 code. It fails by returning a 200 OK, in 3.4 seconds, with a wrong answer, built from an outdated document, after a tool call that silently returned an empty list.
Classic APM sees the request. It sees none of the four decisions that produced the bad answer.
| Question asked in operations | What classic APM says | What you need to instrument |
|---|---|---|
| Why is this answer wrong? | nothing | context content, retrieved documents, actual prompt |
| Why has the bill doubled? | nothing | input and output tokens, per model and per route |
| Why has it been slow since Tuesday? | overall HTTP latency | breakdown embeddings / search / reranking / inference |
| Did this tool really work? | server 200 code | MCP call result, error.type, size of the response |
| Is quality drifting? | nothing | online evaluation scores, correlated with the trace |
Three distinct planes must be covered and a stack that covers only one leaves the other two in the dark:
- the inference plane: model calls, tokens, latency, time to first token, cost;
- the retrieval plane: vectorization, search, reranking, context assembly;
- the tools plane: MCP calls, process boundary crossing, tool errors.
What follows is a series of deployment playbooks, from the minimal foundation to the production stack, with complete configurations. The building blocks are open source; for Langfuse, only the core is (see section 10).
2. The semantic foundation: OpenTelemetry GenAI
Section titled “2. The semantic foundation: OpenTelemetry GenAI”2.1 State of the standard in October 2026
Section titled “2.1 State of the standard in October 2026”The point that decides how long your stack will last is not the choice of backend, it is the vocabulary. This section is the site’s reference on these conventions: other content links here rather than retelling their history.
- The
gen_ai.*conventions have left the main repository: version 1.42.0 of the semantic conventions (12 June 2026) deprecated thegen_ai.*,openai.*andmcp.*attributes, metrics, events and spans and moved them to theopen-telemetry/semantic-conventions-genairepository. The main repository carried on without them (1.43.0 in July 2026, 1.44.0 in August 2026). - That repository has no tagged release yet (checked on 4 October 2026). Its manifest declares a development schema (
gen-ai-dev/1.42.0-dev) and builds on the general conventions 1.44.0. So you cannot pin a published version: you pin a commit. - No
gen_ai.*attribute, span, metric or event is stable. All carry the Development status. The general attributes they reuse, such aserror.typeornetwork.transport, are stable. - The
mcp.*attributes, introduced in 1.39.0, have followed the same path: deprecated in the main registry, they now live in the GenAI repository.
The renames already absorbed set the pace:
| Version | Change | Consequence in operations |
|---|---|---|
| 1.37.0 (August 2025) | gen_ai.system → gen_ai.provider.name; per-message events (gen_ai.user.message, gen_ai.choice…) give way to the gen_ai.system_instructions, gen_ai.input.messages and gen_ai.output.messages attributes and the gen_ai.client.inference.operation.details event | dual reading mandatory during the transition; content masking rules to rewrite |
| 1.38.0 (October 2025) | gen_ai.evaluation.result event | quality scores enter the bus |
| 1.39.0 (January 2026) | MCP conventions (mcp.*) | MCP tool calls become instrumentable as standard |
| 1.40.0 (February 2026) | retrieval spans and cache attributes | RAG becomes instrumentable as standard |
| 1.41.0 (April 2026) | invoke_agent split into client and internal variants; tool name in the execute_tool span name; time-to-first-chunk metrics | agent dashboards break |
| 1.42.0 (June 2026) | exit from the main repository | no more published schema version to cite |
| GenAI repository, main branch, unreleased | gen_ai.client.token.usage and the gen_ai.token.type attribute replaced by gen_ai.client.inference.usage.* counters (broken down by gen_ai.token.modality) and two per-operation histograms; gen_ai.usage.cache_creation.input_tokens renamed gen_ai.usage.cache_write.input_tokens | the cost rules in section 5.4 read the old name: translate in the Collector the day your instrumentation changes |
2.2 The useful vocabulary
Section titled “2.2 The useful vocabulary”gen_ai.operation.name carries the nature of the operation. The values that matter to you in practice:
| Operation | Span name | What it covers |
|---|---|---|
chat | chat {gen_ai.request.model} | conversational inference call |
embeddings | embeddings {gen_ai.request.model} | vectorization of a query or a document |
retrieval | retrieval {gen_ai.data_source.id} | search in a vector database or an index |
execute_tool | execute_tool {gen_ai.tool.name} | execution of a tool by the agent |
invoke_agent | invoke_agent {gen_ai.agent.name} | complete agent turn |
The attributes to treat as the hard core:
INFERENCE gen_ai.provider.name openai | anthropic | mistral_ai | aws.bedrock ... (ollama, qdrant: custom values) gen_ai.request.model requested model gen_ai.response.model model actually served (may differ) gen_ai.usage.input_tokens integer gen_ai.usage.output_tokens integer gen_ai.response.finish_reasons array: stop | length | tool_call | error ... gen_ai.conversation.id conversation identifier
RETRIEVAL gen_ai.data_source.id source identifier gen_ai.retrieval.top_k number of documents requested gen_ai.retrieval.query.text opt-in, sensitive content gen_ai.retrieval.documents opt-in, sensitive content
MCP mcp.method.name tools/call | resources/read | server/discover mcp.protocol.version 2026-07-28 ... mcp.session.id versions before 2026-07-28, HTTP only network.transport pipe (stdio) | tcp (HTTP) error.type set only on failureFour clarifications, checked against the conventions repository on 4 October 2026:
gen_ai.provider.namehas a list of well-known values (openai,anthropic,mistral_ai,aws.bedrock,azure.ai.openai,gcp.vertex_ai,cohere,deepseek,groq…): if one applies, it must be used; otherwise a custom value is allowed.ollamais not one of them: it is a legitimate custom value, likeqdranton the retrieval span in playbook 3. Watch out for Mistral AI, whose value ismistral_ai, notmistral.gen_ai.response.finish_reasonsis an array, one reason per returned choice, with no closed list in the registry (examples:stop,length,error). The output messages schema listsstop,length,content_filter,tool_call,compactionanderror: it is indeedtool_call, singular, not the rawtool_callsvalue some providers return. A reason that was expected but not received (failed generation, interrupted stream) must be reported aserror.mcp.method.namefollows the same well-known values rule. The registry still reflects the methods that predate the MCP 2026-07-28 specification (initializeis among them):server/discoveris a custom value there for now.- Metrics: two cover most of the day-to-day need, in my view.
gen_ai.client.operation.duration(histogram, in seconds) gives latency per operation. For tokens, until they moved, the conventions described thegen_ai.client.token.usagehistogram, broken down bygen_ai.token.type; the main branch of the GenAI repository replaces it with thegen_ai.client.inference.usage.input_tokensandgen_ai.client.inference.usage.output_tokenscounters. Check which one your instrumentation emits: the rules in section 5.4 are written ongen_ai.client.token.usage, which is exactly the case where translation in the Collector protects you. For time to first chunk when streaming,gen_ai.client.operation.time_to_first_chunkhas existed since 1.41.0.
Enabling it on the SDK side:
export OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimentalNot all SDKs honor this flag in the same way. Check what actually comes out before building dashboards on it: the acceptance command is given in playbook 1.
2.3 Content capture: the rule
Section titled “2.3 Content capture: the rule”Prompts, system instructions and responses are the most sensitive data in the stack. The conventions are explicit: instrumentations should not capture this content by default and should offer an explicit opt-in. They describe three modes:
- by default, no capture;
- on spans, in the
gen_ai.system_instructions,gen_ai.input.messagesandgen_ai.output.messagesattributes, all at the Opt-In level: suited when volume stays manageable and either privacy regulations do not apply or the telemetry storage complies with them, for example in pre-production; - in external storage, with only a reference on the span: the mode recommended in production, because it allows separate access controls.
The same content can also travel in the gen_ai.client.inference.operation.details event, also Opt-In, which stores it independently from traces. The gen_ai.retrieval.query.text, gen_ai.retrieval.documents and gen_ai.tool.definitions attributes are Opt-In too.
The rule adopted across the whole site, whose wording other content reuses: content capture is disabled by default in production; it is enabled explicitly, per environment, with its own retention; preferably outside the indexed span attributes (dedicated event or log, or referenced external storage), with masking at the Collector; and its absence is verified by an actual search in each backend. The implementation, application barrier and Collector barrier, is detailed in section 9.
2.4 Application-specific attributes
Section titled “2.4 Application-specific attributes”The gen_ai.* namespace is reserved for the conventions. Attributes specific to an application use a namespace of its own: in these playbooks llm.* and rag.*, elsewhere on the site mttl.* or app.*, never gen_ai.*. A home-grown attribute placed under gen_ai.* would collide with the next version of the convention.
3. The target architecture
Section titled “3. The target architecture”One bus, several consumers. This is, in my view, the most structuring architecture decision.
+---------------------------------------------------------------+| APPLICATION: agent, RAG chain, MCP client || OpenTelemetry SDK + GenAI instrumentation |+---------------------------------------------------------------+ v OTLP 4317 / 4318+---------------------------------------------------------------+| OTEL COLLECTOR: gateway || normalization > redaction > sampling > derivation |+---------------------------------------------------------------+ v v v v+---------------------------------------------------------------+| Tempo/Jaeger Prometheus/VM Loki/OpenSearch Langfuse || traces metrics logs LLM + evals |+---------------------------------------------------------------+ v+---------------------------------------------------------------+| GRAFANA: operations, cost, quality |+---------------------------------------------------------------+The non-negotiable principle: the application emits only OTLP. It does not know that Langfuse, Tempo or Prometheus exist. The day you change LLM backend, you do not redeploy forty services: you change the exporters section of the Collector.
The symmetrical mistake, to be avoided, is to run two instrumentation SDKs side by side, the LLM observability vendor’s and OpenTelemetry’s. You get two disjoint trace trees, double ingestion billing and trace identifiers that do not correlate. One SDK, one gateway.
4. Playbook 1: Minimal foundation in half a day
Section titled “4. Playbook 1: Minimal foundation in half a day”Goal: see the first complete LLM trace, with tokens and cost, without deciding anything final.
4.1 Components
Section titled “4.1 Components”| Role | Building block | License |
|---|---|---|
| Instrumentation | OpenLIT or OpenLLMetry (Traceloop) | Apache-2.0 |
| Gateway | OpenTelemetry Collector contrib | Apache-2.0 |
| LLM backend | Langfuse (self-hosted) | MIT (core) |
4.2 Deploying Langfuse
Section titled “4.2 Deploying Langfuse”Langfuse v3 is not a single binary: it needs PostgreSQL, ClickHouse, Redis and S3-compatible object storage. Do not rebuild this composition by hand, start from the maintained one:
git clone --depth 1 https://github.com/langfuse/langfuse.gitcd langfuse && docker compose up -dThe interface is on http://localhost:3000. Create a project, retrieve the pk-lf-... / sk-lf-... key pair, then build the OTLP authentication header:
export LANGFUSE_BASIC_AUTH=$(printf '%s:%s' "$LANGFUSE_PUBLIC_KEY" "$LANGFUSE_SECRET_KEY" | base64 -w0)4.3 The Collector
Section titled “4.3 The Collector”Langfuse exposes a native OTLP endpoint on /api/public/otel, over HTTP only, no gRPC. The Collector absorbs this constraint: the application can keep speaking gRPC if it wishes.
otel-collector.yaml:
receivers: otlp: protocols: grpc: endpoint: 0.0.0.0:4317 http: endpoint: 0.0.0.0:4318
processors: memory_limiter: check_interval: 1s limit_percentage: 80 spike_limit_percentage: 20 batch: timeout: 2s send_batch_size: 512
# Dual reading: gen_ai.system (< 1.37) -> gen_ai.provider.name transform/semconv: error_mode: ignore trace_statements: - set(span.attributes["gen_ai.provider.name"], span.attributes["gen_ai.system"]) where span.attributes["gen_ai.provider.name"] == nil and span.attributes["gen_ai.system"] != nil
exporters: otlphttp/langfuse: endpoint: http://langfuse-web:3000/api/public/otel headers: Authorization: "Basic ${env:LANGFUSE_BASIC_AUTH}" debug: verbosity: detailed
service: pipelines: traces: receivers: [otlp] processors: [memory_limiter, transform/semconv, batch] exporters: [otlphttp/langfuse, debug]4.4 Instrumenting the application
Section titled “4.4 Instrumenting the application”pip install openlitimport openlit
openlit.init( application_name="assistant-support", environment="lab", otlp_endpoint="http://otel-collector:4318", capture_message_content=False, # read section 9 before switching to True)Environment variables, to be set at deployment level rather than in the code:
OTEL_SERVICE_NAME=assistant-supportOTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318OTEL_EXPORTER_OTLP_PROTOCOL=http/protobufOTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimentalOTEL_RESOURCE_ATTRIBUTES=deployment.environment.name=lab,service.version=1.4.24.5 Acceptance
Section titled “4.5 Acceptance”Do not validate on “I can see something in the interface”. Validate on the attributes:
docker compose logs otel-collector | grep -E 'gen_ai\.(provider\.name|request\.model|usage\.)' | head -20| Check | Expected |
|---|---|
| One span per model call | name chat <model> |
| Tokens present | gen_ai.usage.input_tokens and output_tokens non-zero |
| Normalized provider | gen_ai.provider.name set, including on older SDKs |
| Single trace | all spans of a request share the same trace_id |
| No content | neither prompt nor completion in the debug output |
If the last point fails, stop there and deal with section 9 before going any further.
5. Playbook 2: Production stack
Section titled “5. Playbook 2: Production stack”The playbook 1 foundation does not hold up in production: it sends 100% of traces, derives no metrics and has only one consumer.
5.1 Two-tier topology
Section titled “5.1 Two-tier topology”Tail sampling requires that all spans of a trace land on the same Collector instance. With several replicas, this calls for a first tier that distributes by trace_id.
Whichever agent receives a span, the routing key sends it to the gateway that holds its trace:
flowchart TB
APP["Applications<br/>OTLP 4317 / 4318"]
subgraph E1["Tier 1: agents"]
A1["Agent 1<br/>memory_limiter, batch,<br/>load_balancing"]
A2["Agent 2<br/>memory_limiter, batch,<br/>load_balancing"]
end
subgraph E2["Tier 2: gateways"]
G1["Gateway A<br/>tail_sampling, redaction,<br/>connectors"]
G2["Gateway B<br/>tail_sampling, redaction,<br/>connectors"]
end
APP --> A1
APP --> A2
A1 -->|"trace X"| G1
A2 -->|"trace X"| G1
A1 -->|"trace Y"| G2
A2 -->|"trace Y"| G2
G1 --> TP["Tempo<br/>traces"]
G1 --> VM["VictoriaMetrics<br/>metrics"]
G1 --> LF["Langfuse<br/>LLM"]
G2 --> TP
G2 --> VM
G2 --> LF
Tier 1, the essentials:
exporters: # formerly "loadbalancing", name kept as a deprecated alias load_balancing: routing_key: traceID protocol: otlp: tls: insecure: true resolver: dns: hostname: otel-gateway-headless.observability.svc.cluster.local port: 43175.2 Tail sampling
Section titled “5.2 Tail sampling”In my view, random sampling is a bad choice here: it throws away as many incidents as nominal traffic. The decision must be made once the trace is complete.
processors: tail_sampling: decision_wait: 30s num_traces: 100000 expected_new_traces_per_sec: 500 policies: - name: erreurs type: status_code status_code: status_codes: [ERROR]
- name: outil-en-echec type: string_attribute string_attribute: key: error.type values: [".+"] enabled_regex_matching: true
- name: latence type: latency latency: threshold_ms: 8000
- name: generation-anormalement-longue type: numeric_attribute numeric_attribute: key: gen_ai.usage.output_tokens min_value: 2000
# min_value: 0 and max_value: 0 would be ignored (0 means "not set" # for numeric_attribute): we use an OTTL condition instead. - name: recuperation-vide type: ottl_condition ottl_condition: error_mode: ignore span: - span.attributes["rag.documents.count"] == 0
- name: nominal type: probabilistic probabilistic: sampling_percentage: 5The policies are combined with a logical OR: a trace selected by any of them is kept. The intended result: 100% of the pathological, 5% of the rest.
5.3 Deriving metrics from spans
Section titled “5.3 Deriving metrics from spans”Rather than instrumenting twice, let the Collector build the RED metrics from the spans:
processors: # gen_ai.response.finish_reasons is an array (one reason per choice): # we extract a string from it, into an attribute of the internal model. transform/finish_reason: error_mode: ignore trace_statements: - set(span.attributes["llm.finish_reason"], span.attributes["gen_ai.response.finish_reasons"][0]) where IsList(span.attributes["gen_ai.response.finish_reasons"]) - set(span.attributes["llm.finish_reason"], span.attributes["gen_ai.response.finish_reasons"]) where IsString(span.attributes["gen_ai.response.finish_reasons"])
connectors: spanmetrics: histogram: explicit: buckets: [100ms, 250ms, 500ms, 1s, 2s, 5s, 10s, 30s, 60s] dimensions: - name: gen_ai.provider.name - name: gen_ai.request.model - name: gen_ai.operation.name - name: mcp.method.name - name: deployment.environment.name - name: llm.finish_reasonThe connector copies the attribute value as is: an array would stay an array and the label obtained after translation to Prometheus would be a serialized form (such as ["length"]) that the ="length" filter would not find. Hence the transform/finish_reason processor, which extracts a string from it. It must appear in the traces pipeline before the spanmetrics exporter. It keeps only the reason of the first choice, which is enough for calls with a single completion (n=1), the common case. The attribute produced belongs to the internal model (llm.*) and not to gen_ai.*, following the rule in section 2.4.
Watch the cardinality. Each dimension multiplies the number of active series. The six above are bounded. gen_ai.conversation.id, a user identifier or the name of a retrieved document are not: these values belong in traces, never in metric labels. The “Cost is decided at the Collector” simulator puts figures on the effect of these choices.
5.4 Translating tokens into euros
Section titled “5.4 Translating tokens into euros”Prometheus does not know your rates. Expose them as a static series, through the textfile collector of node_exporter or a twenty-line exporter. Prices change fast and vary by contract: this article does not give any; use your provider’s dated price list, converted into your currency at a date you record. Replace each <...> with a numeric value in euros per token:
# To be filled in from your provider's dated price list (euros per token)genai_prix_euro_par_token{model="grand-modele",token_type="input"} <prix_entree_grand>genai_prix_euro_par_token{model="grand-modele",token_type="output"} <prix_sortie_grand>genai_prix_euro_par_token{model="petit-modele",token_type="input"} <prix_entree_petit>genai_prix_euro_par_token{model="petit-modele",token_type="output"} <prix_sortie_petit>As long as this file does not exist, the join below returns nothing: genai:cout_euro_par_heure is not produced, but genai:jetons:rate5m is and you can alert on tokens with the same rule comparing to the seven-day baseline.
Then the recording rules:
groups: - name: genai-cout interval: 60s rules: - record: genai:jetons:rate5m expr: | sum by (service_name, gen_ai_request_model, gen_ai_token_type) ( rate(gen_ai_client_token_usage_sum[5m]) )
- record: genai:cout_euro_par_heure expr: | sum by (service_name, gen_ai_request_model) ( genai:jetons:rate5m * 3600 * on (gen_ai_request_model, gen_ai_token_type) group_left() label_replace( label_replace(genai_prix_euro_par_token, "gen_ai_request_model", "$1", "model", "(.*)"), "gen_ai_token_type", "$1", "token_type", "(.*)") )The alerts that are actually useful:
- alert: CoutHoraireLLMAnormal expr: | genai:cout_euro_par_heure > 3 * avg_over_time(genai:cout_euro_par_heure[7d] offset 1h) for: 15m annotations: summary: "Cout LLM x3 vs reference 7 jours"
- alert: TroncatureDeGeneration expr: | sum(rate(traces_span_metrics_calls_total{ llm_finish_reason="length"}[10m])) / sum(rate(traces_span_metrics_calls_total{ gen_ai_operation_name="chat"}[10m])) > 0.05 for: 10m annotations: summary: "Plus de 5% des generations coupees par la limite de jetons"The second alert relies on the llm.finish_reason dimension added in 5.3. In my view, it is rarely set up even though it can explain part of the “incomplete answers” reported by users.
6. Playbook 3: Instrumenting RAG
Section titled “6. Playbook 3: Instrumenting RAG”6.1 What really breaks in a RAG
Section titled “6.1 What really breaks in a RAG”| Failure | User symptom | Signal that reveals it | Signal that does not reveal it |
|---|---|---|---|
| Empty or off-topic retrieval | “it makes things up” | rag.documents.count, top-1 score | latency, HTTP code |
| Context retrieved but ignored | “it misses the point” | ratio of cited / retrieved documents | everything else |
| Stale index | “it gives the old procedure” | rag.index.version, index age | model quality |
| Truncated context | partial answer | rag.context.truncated, finish_reasons=length | retrieval scores |
| Embedder drift | slow, diffuse degradation | score distribution over time | an isolated trace |
The last two rows are invisible on a single trace. They only appear in the distribution, so in metrics, not in traces.
6.2 Instrumentation
Section titled “6.2 Instrumentation”The chain splits into three spans, not one.
from opentelemetry import trace
tracer = trace.get_tracer("rag.pipeline")
INDEX = "kb-support-fr"INDEX_VERSION = "v7"MODELE = "grand-modele" # illustrative name
def repondre(question: str) -> str: with tracer.start_as_current_span("embeddings text-embedding-3-large") as sp: sp.set_attribute("gen_ai.operation.name", "embeddings") sp.set_attribute("gen_ai.provider.name", "openai") sp.set_attribute("gen_ai.request.model", "text-embedding-3-large") vecteur = embedder.encode(question)
with tracer.start_as_current_span(f"retrieval {INDEX}") as sp: sp.set_attribute("gen_ai.operation.name", "retrieval") sp.set_attribute("gen_ai.data_source.id", INDEX) sp.set_attribute("gen_ai.provider.name", "qdrant") sp.set_attribute("gen_ai.retrieval.top_k", 8) sp.set_attribute("rag.index.version", INDEX_VERSION)
docs = store.search(vecteur, k=8)
sp.set_attribute("rag.documents.count", len(docs)) sp.set_attribute("rag.score.top", docs[0].score if docs else 0.0) sp.set_attribute("rag.score.min_kept", docs[-1].score if docs else 0.0) # identifiers, never the content sp.set_attribute("rag.document.ids", [d.id for d in docs][:8])
with tracer.start_as_current_span(f"chat {MODELE}") as sp: contexte, tronque = assembler(docs, budget=12_000) sp.set_attribute("rag.context.chars", len(contexte)) sp.set_attribute("rag.context.truncated", tronque) reponse = client.chat(question, contexte) sp.set_attribute("rag.documents.cited", compter_citations(reponse, docs)) return reponseThree choices deserve an explanation.
rag.index.version on every retrieval span. Without this attribute, a failed reindexing cannot be detected after the fact: you will see a drop in quality without being able to date it or attribute it. With it, a single query is enough to confirm the correlation.
Document identifiers, not their content. An identifier is bounded, not personal and enough to replay the retrieval. The content of the chunk is bulky, often confidential and has no place in a span by default. The gen_ai.retrieval.documents attribute exists in the convention, where it is classified as opt-in; this classification is a warning, not a formality.
rag.documents.cited. In my view, this is the most useful measure in the whole chain. The cited / count ratio measures actual use of the context. A ratio that collapses while retrieval scores stay good signals that the problem has moved from the retriever to the generator, typically after a change of model or system prompt.
6.3 The minimal dashboard
Section titled “6.3 The minimal dashboard”Span attributes do not become metrics on their own. The rag_* series below are application metrics that you emit yourself, with the OpenTelemetry SDK, alongside the span attributes:
from opentelemetry import metrics
meter = metrics.get_meter("rag.pipeline")
documents = meter.create_histogram( "rag.documents.count", explicit_bucket_boundaries_advisory=[0, 1, 2, 4, 8])score_top = meter.create_histogram( "rag.score.top", explicit_bucket_boundaries_advisory=[0.2, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0])assemblages = meter.create_counter("rag.context.truncated")
# in repondre(), after the search and then after context assemblyetiquettes = {"rag.index.version": INDEX_VERSION}documents.record(len(docs), etiquettes)score_top.record(docs[0].score if docs else 0.0, etiquettes)assemblages.add(1, {**etiquettes, "truncated": str(tronque).lower()})Once exported over OTLP to a Prometheus-compatible backend, with the usual name translation (dots replaced by underscores, _total suffix for counters, no declared unit), they give the following queries:
sum(rate(rag_documents_count_bucket{le=~"0(\\.0)?"}[5m])) / sum(rate(rag_documents_count_count[5m])) -> empty retrieval rate
histogram_quantile(0.5, sum by (le, rag_index_version) ( rate(rag_score_top_bucket[30m]))) -> median top-1 score
sum(rate(rag_context_truncated_total{truncated="true"}[15m])) / sum(rate(rag_context_truncated_total[15m])) -> truncation rateA drop in the median top-1 score that coincides with a change in rag_index_version is a diagnosis, not a hypothesis.
7. Playbook 4: Instrumenting MCP
Section titled “7. Playbook 4: Instrumenting MCP”7.1 The process boundary problem
Section titled “7.1 The process boundary problem”An MCP server over stdio is a separate process that communicates in JSON-RPC over pipes. There are no HTTP headers, so no traceparent where OpenTelemetry propagators usually look for it. Without specific handling, every tool call produces an orphan trace: you have the latency, you have lost the link with the conversation that triggered it.
The answer is standardized. SEP-414 of the MCP protocol, integrated into the 2026-07-28 specification, documents the use of params._meta as the carrier of the W3C context, with an explicit exception to the DNS prefixing rule for keys.
{ "jsonrpc": "2.0", "id": 2, "method": "tools/call", "params": { "name": "get_weather", "arguments": { "location": "New York" }, "_meta": { "traceparent": "00-0af7651916cd43dd8448eb211c80319c-00f067aa0ba902b7-01" } }}The keys are traceparent, tracestate and baggage, without a prefix. An implementation that wrote io.modelcontextprotocol.traceparent would break correlation with the whole existing ecosystem.
MCP CLIENT: agent MCP SERVER: tool+-------------------------------+ +-------------------------------+| span kind = CLIENT | | span kind = SERVER || name: tools/call get_weather | | name: tools/call get_weather || mcp.method.name = tools/call | | mcp.method.name = tools/call || mcp.protocol.version | | mcp.protocol.version || = 2026-07-28 | | = 2026-07-28 || network.transport = pipe | | network.transport = pipe |+-------------------------------+ +-------------------------------+ JSON-RPC: params._meta.traceparent = 00-<trace>-<span>-01 same trace_id on both sides, parent-child relationship preservedThe same mechanism, unfolded over time across a complete agent turn:
sequenceDiagram participant U as User participant A as Agent, MCP client participant L as LLM participant S as MCP server U->>A: question A->>L: chat, CLIENT span L-->>A: finish_reason tool_call A->>S: tools/call with params._meta.traceparent Note over A,S: same trace_id, the SERVER span is a child of the CLIENT span S-->>A: result, or error with error.type A->>L: chat with the tool result L-->>A: final answer A-->>U: answer
7.2 Client side
Section titled “7.2 Client side”from opentelemetry import tracefrom opentelemetry.propagate import injectfrom opentelemetry.trace import SpanKind
tracer = trace.get_tracer("mcp")
VERSION_MCP = "2026-07-28"
def appeler_outil(message: dict, serveur: str): methode = message["method"] with tracer.start_as_current_span( f"{methode} {serveur}", kind=SpanKind.CLIENT ) as sp: sp.set_attribute("mcp.method.name", methode) sp.set_attribute("mcp.protocol.version", VERSION_MCP) sp.set_attribute("network.transport", "pipe")
meta = message.setdefault("params", {}).setdefault("_meta", {}) # since 2026-07-28, every request carries its protocol version # and the client capabilities (two mandatory keys) meta["io.modelcontextprotocol/protocolVersion"] = VERSION_MCP meta.setdefault("io.modelcontextprotocol/clientCapabilities", {}) inject(meta) # traceparent WITHOUT DNS prefix, SEP-414
reponse = transport.send(message) if "error" in reponse: sp.set_attribute("error.type", str(reponse["error"].get("code"))) sp.set_status(trace.StatusCode.ERROR) return reponse7.3 Server side
Section titled “7.3 Server side”from opentelemetry import tracefrom opentelemetry.propagate import extractfrom opentelemetry.trace import SpanKind
tracer = trace.get_tracer("mcp")
def traiter(message: dict): porteur = (message.get("params") or {}).get("_meta") or {} ctx = extract(porteur) # reads traceparent / tracestate / baggage methode = message["method"]
with tracer.start_as_current_span( f"{methode} {NOM_SERVEUR}", context=ctx, kind=SpanKind.SERVER ) as sp: sp.set_attribute("mcp.method.name", methode) version = porteur.get("io.modelcontextprotocol/protocolVersion") if version: sp.set_attribute("mcp.protocol.version", version) sp.set_attribute("network.transport", "pipe") if methode == "tools/call": sp.set_attribute("gen_ai.tool.name", message["params"]["name"]) return executer(message)If your client or server uses an official MCP SDK, part of this work may already be done upstream and third-party instrumentations also exist. Check before instrumenting by hand: double instrumentation produces duplicate spans and skews every count.
Sessions and handshake. The MCP 2026-07-28 specification makes the protocol stateless: it removes the initialize handshake, protocol sessions and the Mcp-Session-Id header. Every request carries its protocol version and the client capabilities in _meta and a server announces the versions it accepts through the server/discover method. The session identifier never appeared in the JSON-RPC message anyway: in earlier versions, it was an HTTP header of the Streamable HTTP transport, absent over stdio. If you still have to serve clients older than 2026-07-28 over HTTP, read mcp.session.id from that header, not from the message.
7.4 What to monitor on MCP
Section titled “7.4 What to monitor on MCP”| Indicator | Breakdown | Why |
|---|---|---|
| Failure rate per tool | error.type by gen_ai.tool.name | a tool that fails 30% of the time degrades the agent with no visible error |
| p95 latency per tool | mcp.method.name="tools/call" | a slow tool is multiplied by the number of agent turns |
| Calls per agent turn | MCP spans under invoke_agent | detection of tool call loops |
| Tools called outside the list | gen_ai.tool.name outside the allowlist | prompt drift or indirect injection |
| Requests rejected for protocol version | UnsupportedProtocolVersion errors by mcp.protocol.version | clients or servers not yet migrated to 2026-07-28 |
The fourth row deserves a word. An MCP server is an execution surface reachable by the content the model processes. MCP observability is therefore not only a performance topic: the list of tools actually called, compared with the list of tools expected for a given route, is, in my view, an important security signal. It is also one of the few places where an indirect injection can become visible.
8. Playbook 5: Closing the quality loop
Section titled “8. Playbook 5: Closing the quality loop”Without a quality measure, the stack tells you everything is fine while relevance collapses.
The two stages of the loop, offline before deployment and online after, answer each other:
flowchart TB
CH["Change<br/>prompt, model,<br/>chunking, index"] --> CI{"Offline in CI<br/>reference set"}
CI -->|"below threshold"| BL["Deployment<br/>blocked"]
CI -->|"thresholds met"| PR["Production"]
PR --> EC["1 to 3% sample<br/>pinned judge"]
EC --> EV["gen_ai.evaluation.result<br/>correlated with the span"]
EV --> AL{"Drift<br/>alert"}
AL -->|"yes"| DG["Diagnosis<br/>index, prompt, model"]
DG --> CH
DG -.->|"new cases"| CI
8.1 Offline, in continuous integration
Section titled “8.1 Offline, in continuous integration”A versioned reference set, run on every change of prompt, model, chunking or index.
pip install "ragas==0.4.3"import mathimport statistics
from ragas import evaluate# Legacy import, accepted by evaluate() in 0.4.3 with a deprecation# warning (the new metrics live in ragas.metrics.collections).from ragas.metrics import ( faithfulness, answer_relevancy, context_precision, context_recall,)
resultat = evaluate( dataset=jeu_de_reference, # question, context, answer, ground truth metrics=[faithfulness, answer_relevancy, context_precision, context_recall],)
SEUILS = {"faithfulness": 0.85, "context_recall": 0.80}for metrique, seuil in SEUILS.items(): # resultat[metrique] is the list of per-example scores: we compare the mean scores = [s for s in resultat[metrique] if s is not None and not math.isnan(s)] moyenne = statistics.fmean(scores) if scores else 0.0 if moyenne < seuil: raise SystemExit(f"REGRESSION {metrique}: {moyenne:.3f} < {seuil}")The 0.85 and 0.80 thresholds are examples, to be calibrated on your own measurements. The four metrics do not measure the same thing and the diagnosis depends on which one drops:
| Metric dropping | Component at fault |
|---|---|
context_recall | chunking, embedder, top_k too low |
context_precision | no reranking, score threshold too permissive |
faithfulness | system prompt, model, truncated context |
answer_relevancy | system prompt, rephrasing of the question |
8.2 Online, on a production sample
Section titled “8.2 Online, on a production sample”The reference set does not contain the questions your users will ask tomorrow. You need to judge a sample of real traffic. The site’s common principle, detailed in the guide, section 5.7: what counts is the number of judgments per segment and per period, not the percentage. A rate of 1 to 3% is an example starting point, to be adjusted to get enough judgments per alert window; since each judgment is itself a model call, cost quickly limits this rate.
The result is written to the bus as a gen_ai.evaluation.result event, introduced in 1.38.0, correlated with the original span. An event does not become a metric on its own: the evaluation service also emits an application histogram, for example meter.create_histogram("genai.evaluation.score") recorded with the evaluation_name attribute, whose bucket boundaries include the quality threshold (for example explicit_bucket_boundaries_advisory=[0.5, 0.6, 0.7, 0.8, 0.9, 1.0]). It thus becomes an alertable metric.
The alert is on a proportion, not an average: the share of judged responses whose score is above the threshold, compared with a target. An average can stay flat while a growing fraction of responses falls below the threshold.
- alert: SLOQualiteFidelite expr: | ( 1 - sum(rate(genai_evaluation_score_bucket{evaluation_name="faithfulness", le="0.8"}[6h])) / sum(rate(genai_evaluation_score_count{evaluation_name="faithfulness"}[6h])) ) < 0.95 and sum(increase(genai_evaluation_score_count{evaluation_name="faithfulness"}[6h])) >= 100 for: 30m annotations: summary: "Under 95% of judged responses above 0.8 faithfulness over 6h"The threshold (0.8), the target (95%) and the minimum of 100 judgments are examples. This rule still compares a point value with the target: the GenAI method, part IV compares the lower bound of the confidence interval (Wilson bound), sizes the sample and derives error budget and multi-window burn rate from it. It is the reference for quality SLOs.
Three precautions, in order of importance:
- Pin the judge model version. A judge that changes version shifts all your scores and you will spend a week looking for an application regression that does not exist.
- Do not judge with the model being judged. A judge model tends to favor its own outputs: this self-preference bias was measured by Panickssery, Bowman and Feng (NeurIPS 2024) on summarization tasks (XSUM, CNN/DailyMail), with a magnitude that varies across models.
- Budget for the judge. For example, at 3% of 500,000 requests per day, that makes 15,000 daily calls, which will show up in the playbook 2 cost dashboard. Label them with a distinct
service.nameso as not to confuse them with application traffic.
9. Confidentiality: do not turn observability into a leak
Section titled “9. Confidentiality: do not turn observability into a leak”This is the point where an LLM observability stack becomes an incident.
Prompts and completions contain, by construction, what the user wrote: personal data, health data, contractual elements, sometimes a secret pasted into a chat window. The documents retrieved by RAG contain the internal knowledge base, including what the user had no right to see. All of this is replicated, indexed and retained.
| Data | Risk | Where to handle it |
|---|---|---|
| User prompt | personal data, secrets | capture disabled by default, masking at the Collector |
| Completion | restitution of context data | same treatment, short retention |
| Retrieved documents | knowledge base leak, bypassing access rights | identifiers only, never the content |
| MCP call arguments | identifiers, tokens, paths | argument allowlist, no bulk capture |
gen_ai.conversation.id | re-identification by cross-referencing | pseudonymize, never as a metric label |
The rule, set out in section 2.3: content is disabled by default and enabled explicitly, per environment, with its own retention. The flag exists on the SDK side (capture_message_content, or an OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT variable depending on the implementation; the name still varies, check yours). It is not enough: a flag gets set back to true by accident, in the middle of a debugging session. The second barrier is in the Collector and that one cannot be bypassed from an application.
processors: transform/confidentialite: error_mode: ignore trace_statements: # 1. Masking of the most common patterns - replace_pattern(span.attributes["gen_ai.input.messages"], "[\\w.+-]+@[\\w-]+\\.[\\w.]{2,}", "[courriel]") - replace_pattern(span.attributes["gen_ai.input.messages"], "\\b(?:\\d[ -]?){13,19}\\b", "[carte]") - replace_pattern(span.attributes["gen_ai.output.messages"], "[\\w.+-]+@[\\w-]+\\.[\\w.]{2,}", "[courriel]")
# 2. Deletion outside authorized environments - delete_key(span.attributes, "gen_ai.input.messages") where resource.attributes["deployment.environment.name"] == "production" - delete_key(span.attributes, "gen_ai.output.messages") where resource.attributes["deployment.environment.name"] == "production" - delete_key(span.attributes, "gen_ai.retrieval.documents")Depending on the version of your instrumentation, the content may arrive not as a span attribute but as a log record, the GenAI events. replace_pattern only acts on a string value: if your instrumentation emits messages in structured form, masking does not apply and only deletion protects you. Finally, duplicate the same rules in log_statements, replacing the span. prefix with log.: a clean traces pipeline paired with an unfiltered logs pipeline is, in my view, the easiest trap to miss in this section.
10. Choosing the building blocks: the licensing question
Section titled “10. Choosing the building blocks: the licensing question”The tool landscape and its licenses have a single home on the site: chapter 6 of the guide. Only the points the playbooks depend on remain here.
| Playbook building block | License | Point of attention |
|---|---|---|
| OpenTelemetry Collector, Jaeger, Prometheus, OpenLIT, OpenLLMetry, Ragas | Apache-2.0 | none for internal use |
| VictoriaMetrics | Apache-2.0 (community edition, cluster version included) | downsampling, multiple retentions, anomaly detection in the enterprise edition |
| Grafana, Tempo, Loki | AGPL-3.0 | modifying the component and opening it to users over the network, even without distributing it, obliges you to offer them the modified source code |
| Langfuse (playbook 1) | MIT for the core | some features under a commercial license; owned by ClickHouse since January 2026 (ClickHouse announcement) |
| Arize Phoenix (kit of the labs) | Elastic License 2.0 (source available) | not open source in the OSI sense: prohibited from being offered as a managed service to third parties; no effect for internal use |
To choose the building block for the LLM plane, apply the same three criteria to every option: the license, the ability to self-host and the ability to receive OTLP with the gen_ai.* conventions without a proprietary SDK. In every case, keep your existing stack for the rest: you do not duplicate Tempo and Prometheus, you add a consumer to the bus.
11. Acceptance testing
Section titled “11. Acceptance testing”The criterion I recommend keeping above all others: start from a user complaint and trace it back to the cause in less than five minutes, without access to the production machine.
- A user request produces a single trace, from the HTTP entry point up to and including the MCP server.
- Input and output tokens are present on 100% of inference spans.
- Hourly cost per service and per model is displayed and alerted on deviation from the weekly baseline.
- Retrieval spans carry
rag.documents.count, the top-1 score and the index version. - The ratio of cited documents to retrieved documents is measured.
- MCP client and server spans share the same
trace_idviaparams._meta.traceparent. - The failure rate is visible per MCP tool, not only in aggregate.
- The rate of generations cut off by the token limit is tracked.
- Prompts and completions are absent from production backends, verified by an actual search.
- An online quality score exists, fed by a sample, with a pinned judge model.
- Sampling keeps 100% of errors and slow traces.
- No unbounded attribute is used as a metric label.
- Switching from one backend to another requires no application change.
If the last box is not ticked, you have not deployed an observability stack: you have deployed a proprietary client with extra steps.
12. Conclusion
Section titled “12. Conclusion”Three ideas to remember, in this order.
Vocabulary comes before tooling. The GenAI conventions are on the move: repository relocated, no stable attribute, regular renames. Instrument towards OpenTelemetry, but build your dashboards and alerts on an internal model that you control and absorb the changes in the Collector.
RAG and MCP are the two planes that stacks forget. Inference is instrumented by default by every library. Retrieval and tool calls are not and they are, in my view, what produces a good share of the bad answers. A retrieval score, an index version and a traceparent in _meta: three elements, half a day of work as an order of magnitude and a large part of incidents becomes diagnosable.
Observing an LLM system is processing personal data. A complete trace contains what the user wrote and what the internal knowledge base contained. Content capture is therefore decided as a compliance measure, with one barrier on the application side and a second on the Collector side, tested by an actual search in the backends. A stack that exposes prompts to the whole operations team is not progress in observability, it is an incident waiting to be declared.
Next in the progression: VictoriaMetrics for LLMs for the production metrics backend, then the GenAI method for quality SLOs, quantified impact and compliance. To go back to the concepts: the guide; to practice on a complete stack: the labs.
Revised on 2 October 2026: prices removed, MCP section updated for the 2026-07-28 specification (no more protocol sessions or initialize handshake, protocol version in _meta), llm.finish_reason dimension added to the spanmetrics connector, rag_* and genai_evaluation_score metrics presented as application metrics to emit, Ragas gate corrected and pinned at 0.4.3, exporter renamed load_balancing, “empty retrieval” policy switched to an OTTL condition, OTTL statements written with prefixed paths, prices made illustrative, VictoriaMetrics license corrected, acquisition of Langfuse by ClickHouse mentioned, choice of the LLM building block reworded around explicit criteria with alternatives, scope of the AGPL clarified, judges’ self-preference bias sourced, estimates requalified as opinions or orders of magnitude, clientCapabilities key added to the MCP client, le filter adapted to Prometheus 3 normalization.
Revised on 4 October 2026: page moved to MDX with the progression box, section 2 made the site’s reference on the OpenTelemetry GenAI conventions (version history dated and checked against the conventions repository, well-known values of gen_ai.provider.name with mistral_ai and ollama presented as a custom value, tool_call singular in gen_ai.response.finish_reasons, ongoing replacement of gen_ai.client.token.usage, execute_tool and invoke_agent span names, content capture rule, namespace for home-grown attributes), quality alert reworded as a proportion with a link to the method, judge sampling tied to the guide’s principle, licensing section reduced to the playbooks’ building blocks with a link to the guide, links to the guide, the labs and the method.