Technical
Part II. Architecture and reference configurations
For: engineers and SREs · architectsPrerequisites: Have read part I; know the OpenTelemetry Collector and the PromQL or MetricsQL language.
5. The unified telemetry plane
Section titled “5. The unified telemetry plane”Do not build a GenAI silo. Emit LLM, agent, RAG and MCP telemetry into the same plane as existing infrastructure and applications. The method is tool-agnostic. VictoriaMetrics, Grafana and Langfuse are one possible instantiation among others, chosen here because it can be self-hosted and relies on open protocols (OTLP, remote write), not a dependency of the method. The same roles can be filled with Prometheus, Mimir or Thanos for metrics, Tempo or Jaeger for traces, Loki for logs, Phoenix (under the Elastic License 2.0, a source-available license) or a commercial service such as LangSmith for evaluation. Langfuse has belonged to ClickHouse since January 2026; its core remains under the MIT license.
flowchart TB APP["Application: LLM / Agent / RAG / MCP clients"] INST["OTel instrumentation<br/>OpenLLMetry, OpenInference, native SDK"] COL["OpenTelemetry Collector<br/>masking, attributes, spanmetrics"] MET["Metrics<br/>VictoriaMetrics<br/>(DCGM, vLLM)"] TRA["Traces<br/>OTLP backend"] LOG["Logs<br/>VictoriaLogs / Loki"] CON["Content opt-in<br/>encrypted object store"] EVAL["Evaluation layer<br/>Langfuse / Phoenix"] GRAF["Grafana<br/>dashboards, alerting, SLOs"] APP --> INST --> COL COL --> MET COL --> TRA COL --> LOG COL --> CON MET --> GRAF TRA --> GRAF LOG --> GRAF TRA --> EVAL CON --> EVAL EVAL --> GRAF
6. Reference collector configuration
Section titled “6. Reference collector configuration”Annotated example, to adapt. Key points: one trace pipeline, one metrics pipeline, masking of sensitive data at the collector, derivation of metrics from spans (spanmetrics) with low-cardinality dimensions.
# otel-collector.yaml (reference excerpt)receivers: otlp: protocols: grpc: { endpoint: 0.0.0.0:4317 } http: { endpoint: 0.0.0.0:4318 }
processors: batch: timeout: 5s send_batch_size: 1000
# Masking: strip or hash sensitive content before any export. redaction: allow_all_keys: true blocked_values: - "[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\\.[A-Za-z]{2,}" # emails - "\\b(?:\\d[ -]*?){13,16}\\b" # card numbers summary: info
# Safety net on traces: if an instrumentation captures content by # mistake, strip it from spans before export. Opt-in content goes to the # encrypted external store, never to the trace backend. # Metrics do not need this filter: spanmetrics only keeps the dimensions # listed below. attributes/clean: actions: - key: gen_ai.input.messages action: delete - key: gen_ai.output.messages action: delete
connectors: # Derive metrics from spans, low-cardinality dimensions. spanmetrics: dimensions: - name: gen_ai.request.model - name: gen_ai.operation.name - name: gen_ai.agent.name - name: error.type
exporters: # Metrics to VictoriaMetrics (remote write endpoint). prometheusremotewrite: endpoint: http://victoriametrics:8428/api/v1/write # Traces to the OTLP backend. otlp/traces: endpoint: tempo:4317 tls: { insecure: true } # Logs to VictoriaLogs or Loki. otlphttp/logs: logs_endpoint: http://victorialogs:9428/insert/opentelemetry/v1/logs
service: pipelines: traces: receivers: [otlp] processors: [attributes/clean, redaction, batch] exporters: [otlp/traces, spanmetrics] metrics: receivers: [otlp, spanmetrics] processors: [batch] exporters: [prometheusremotewrite] logs: receivers: [otlp] processors: [redaction, batch] exporters: [otlphttp/logs]7. MetricsQL recording rules
Section titled “7. MetricsQL recording rules”Examples to set in VictoriaMetrics or a compatible Prometheus. The source metrics are those of the OTel GenAI conventions 1.40 and 1.41 (gen_ai.client.operation.duration, gen_ai.client.token.usage, and gen_ai.client.operation.time_to_first_chunk from 1.41 onwards). The names account for the translation done by the prometheusremotewrite exporter: by default (translation_strategy: UnderscoreEscapingWithSuffixes), it replaces dots with underscores and appends the unit suffix, so gen_ai.client.operation.duration (unit s) becomes gen_ai_client_operation_duration_seconds_bucket, _sum and _count. The {token} unit is an annotation and adds no suffix: gen_ai_client_token_usage_sum stays as is. If you choose UnderscoreEscapingWithoutSuffixes, remove _seconds from the rules. Take prices from your own dated table (annex D).
# rules.yaml (reference excerpt, MetricsQL syntax)groups: - name: genai interval: 30s rules: # p95 latency per model: operation duration. - record: genai:operation_duration:p95 expr: | histogram_quantile(0.95, sum(rate(gen_ai_client_operation_duration_seconds_bucket[5m])) by (le, gen_ai_request_model))
# p95 TTFT per model: time to first chunk, streaming calls only. - record: genai:time_to_first_chunk:p95 expr: | histogram_quantile(0.95, sum(rate(gen_ai_client_operation_time_to_first_chunk_seconds_bucket[5m])) by (le, gen_ai_request_model))
# Error ratio per error type. - record: genai:error_ratio expr: | sum(rate(gen_ai_client_operation_duration_seconds_count{error_type!=""}[5m])) by (error_type) / ignoring(error_type) group_left sum(rate(gen_ai_client_operation_duration_seconds_count[5m]))
# Estimated cost rate (euros per second), via tokens and prices injected as labels. # Divide by the request rate to get a cost per request. # price_in and price_out come from a price recording rule or a relabel. - record: genai:cost_rate:eur_per_second expr: | (sum(rate(gen_ai_client_token_usage_sum{gen_ai_token_type="input"}[5m])) by (gen_ai_request_model) * on(gen_ai_request_model) group_left price_in_eur_per_token) + (sum(rate(gen_ai_client_token_usage_sum{gen_ai_token_type="output"}[5m])) by (gen_ai_request_model) * on(gen_ai_request_model) group_left price_out_eur_per_token)
# Average input tokens per request (prompt bloat detection). - record: genai:input_tokens_per_request:avg expr: | sum(rate(gen_ai_client_token_usage_sum{gen_ai_token_type="input"}[1h])) by (gen_ai_request_model) / sum(rate(gen_ai_client_operation_duration_seconds_count[1h])) by (gen_ai_request_model)The price series used by the cost rule is not magic. You materialize it with a small per-model recording rule, kept in sync with your own dated price table (annex D). Prices change fast and vary by contract: this guide gives none; use your provider’s dated price list. Replace <p_in> and <p_out> with your prices in euros per token:
# <p_in> and <p_out>: fill in from your provider's dated price list.- record: price_in_eur_per_token labels: { gen_ai_request_model: "demo-llm" } expr: vector(<p_in>)- record: price_out_eur_per_token labels: { gen_ai_request_model: "demo-llm" } expr: vector(<p_out>)The cost rule group_left join then reads these series. As long as no price series exists for a model, the join returns nothing and the rule emits no cost for that model: the token rules (genai:input_tokens_per_request:avg) remain available. In production these two series are produced by a job that reads a versioned price table. The public genaiotel lab follows a variant: it computes the cost of each span in the Collector (OTTL) from a price table, then a vmalert rule derives the cost series.
Grafana dashboards: one board per layer, correlated by trace_id, with as main panels p95 latency, error ratio by error.type, the distribution of finish_reasons, input tokens per request over time and cost per request and per feature. Design rule: every panel answers a decision, otherwise it is removed; a panel that represents an SLO has its alert (see the guide, 9.5).
Revised on 2 October 2026: prices removed, metric names fixed (_seconds suffix added by the exporter), TTFT rule added, content filter moved to the traces pipeline, pointer to the public genaiotel lab, mention of the acquisition of Langfuse by ClickHouse. Token metrics are evolving in the semantic-conventions-genai repository (see Part I).
Revised on 4 October 2026: the dashboard design rule points to guide 9.5, the corresponding pitfall having left Part VIII.