Skip to content

TechnicalExpert

Part II. Architecture and reference configurations

For: engineers and SREs · architectsPrerequisites: Have read part I; know the OpenTelemetry Collector and the PromQL or MetricsQL language.

Do not build a GenAI silo. Emit LLM, agent, RAG and MCP telemetry into the same plane as existing infrastructure and applications. The method is tool-agnostic. VictoriaMetrics, Grafana and Langfuse are one possible instantiation among others, chosen here because it can be self-hosted and relies on open protocols (OTLP, remote write), not a dependency of the method. The same roles can be filled with Prometheus, Mimir or Thanos for metrics, Tempo or Jaeger for traces, Loki for logs, Phoenix (under the Elastic License 2.0, a source-available license) or a commercial service such as LangSmith for evaluation. Langfuse has belonged to ClickHouse since January 2026; its core remains under the MIT license.

flowchart TB
  APP["Application: LLM / Agent / RAG / MCP clients"]
  INST["OTel instrumentation<br/>OpenLLMetry, OpenInference, native SDK"]
  COL["OpenTelemetry Collector<br/>masking, attributes, spanmetrics"]
  MET["Metrics<br/>VictoriaMetrics<br/>(DCGM, vLLM)"]
  TRA["Traces<br/>OTLP backend"]
  LOG["Logs<br/>VictoriaLogs / Loki"]
  CON["Content opt-in<br/>encrypted object store"]
  EVAL["Evaluation layer<br/>Langfuse / Phoenix"]
  GRAF["Grafana<br/>dashboards, alerting, SLOs"]
  APP --> INST --> COL
  COL --> MET
  COL --> TRA
  COL --> LOG
  COL --> CON
  MET --> GRAF
  TRA --> GRAF
  LOG --> GRAF
  TRA --> EVAL
  CON --> EVAL
  EVAL --> GRAF

Annotated example, to adapt. Key points: one trace pipeline, one metrics pipeline, masking of sensitive data at the collector, derivation of metrics from spans (spanmetrics) with low-cardinality dimensions.

# otel-collector.yaml (reference excerpt)
receivers:
otlp:
protocols:
grpc: { endpoint: 0.0.0.0:4317 }
http: { endpoint: 0.0.0.0:4318 }
processors:
batch:
timeout: 5s
send_batch_size: 1000
# Masking: strip or hash sensitive content before any export.
redaction:
allow_all_keys: true
blocked_values:
- "[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\\.[A-Za-z]{2,}" # emails
- "\\b(?:\\d[ -]*?){13,16}\\b" # card numbers
summary: info
# Safety net on traces: if an instrumentation captures content by
# mistake, strip it from spans before export. Opt-in content goes to the
# encrypted external store, never to the trace backend.
# Metrics do not need this filter: spanmetrics only keeps the dimensions
# listed below.
attributes/clean:
actions:
- key: gen_ai.input.messages
action: delete
- key: gen_ai.output.messages
action: delete
connectors:
# Derive metrics from spans, low-cardinality dimensions.
spanmetrics:
dimensions:
- name: gen_ai.request.model
- name: gen_ai.operation.name
- name: gen_ai.agent.name
- name: error.type
exporters:
# Metrics to VictoriaMetrics (remote write endpoint).
prometheusremotewrite:
endpoint: http://victoriametrics:8428/api/v1/write
# Traces to the OTLP backend.
otlp/traces:
endpoint: tempo:4317
tls: { insecure: true }
# Logs to VictoriaLogs or Loki.
otlphttp/logs:
logs_endpoint: http://victorialogs:9428/insert/opentelemetry/v1/logs
service:
pipelines:
traces:
receivers: [otlp]
processors: [attributes/clean, redaction, batch]
exporters: [otlp/traces, spanmetrics]
metrics:
receivers: [otlp, spanmetrics]
processors: [batch]
exporters: [prometheusremotewrite]
logs:
receivers: [otlp]
processors: [redaction, batch]
exporters: [otlphttp/logs]

Examples to set in VictoriaMetrics or a compatible Prometheus. The source metrics are those of the OTel GenAI conventions 1.40 and 1.41 (gen_ai.client.operation.duration, gen_ai.client.token.usage, and gen_ai.client.operation.time_to_first_chunk from 1.41 onwards). The names account for the translation done by the prometheusremotewrite exporter: by default (translation_strategy: UnderscoreEscapingWithSuffixes), it replaces dots with underscores and appends the unit suffix, so gen_ai.client.operation.duration (unit s) becomes gen_ai_client_operation_duration_seconds_bucket, _sum and _count. The {token} unit is an annotation and adds no suffix: gen_ai_client_token_usage_sum stays as is. If you choose UnderscoreEscapingWithoutSuffixes, remove _seconds from the rules. Take prices from your own dated table (annex D).

# rules.yaml (reference excerpt, MetricsQL syntax)
groups:
- name: genai
interval: 30s
rules:
# p95 latency per model: operation duration.
- record: genai:operation_duration:p95
expr: |
histogram_quantile(0.95,
sum(rate(gen_ai_client_operation_duration_seconds_bucket[5m])) by (le, gen_ai_request_model))
# p95 TTFT per model: time to first chunk, streaming calls only.
- record: genai:time_to_first_chunk:p95
expr: |
histogram_quantile(0.95,
sum(rate(gen_ai_client_operation_time_to_first_chunk_seconds_bucket[5m])) by (le, gen_ai_request_model))
# Error ratio per error type.
- record: genai:error_ratio
expr: |
sum(rate(gen_ai_client_operation_duration_seconds_count{error_type!=""}[5m])) by (error_type)
/
ignoring(error_type) group_left
sum(rate(gen_ai_client_operation_duration_seconds_count[5m]))
# Estimated cost rate (euros per second), via tokens and prices injected as labels.
# Divide by the request rate to get a cost per request.
# price_in and price_out come from a price recording rule or a relabel.
- record: genai:cost_rate:eur_per_second
expr: |
(sum(rate(gen_ai_client_token_usage_sum{gen_ai_token_type="input"}[5m])) by (gen_ai_request_model) * on(gen_ai_request_model) group_left price_in_eur_per_token)
+
(sum(rate(gen_ai_client_token_usage_sum{gen_ai_token_type="output"}[5m])) by (gen_ai_request_model) * on(gen_ai_request_model) group_left price_out_eur_per_token)
# Average input tokens per request (prompt bloat detection).
- record: genai:input_tokens_per_request:avg
expr: |
sum(rate(gen_ai_client_token_usage_sum{gen_ai_token_type="input"}[1h])) by (gen_ai_request_model)
/
sum(rate(gen_ai_client_operation_duration_seconds_count[1h])) by (gen_ai_request_model)

The price series used by the cost rule is not magic. You materialize it with a small per-model recording rule, kept in sync with your own dated price table (annex D). Prices change fast and vary by contract: this guide gives none; use your provider’s dated price list. Replace <p_in> and <p_out> with your prices in euros per token:

# <p_in> and <p_out>: fill in from your provider's dated price list.
- record: price_in_eur_per_token
labels: { gen_ai_request_model: "demo-llm" }
expr: vector(<p_in>)
- record: price_out_eur_per_token
labels: { gen_ai_request_model: "demo-llm" }
expr: vector(<p_out>)

The cost rule group_left join then reads these series. As long as no price series exists for a model, the join returns nothing and the rule emits no cost for that model: the token rules (genai:input_tokens_per_request:avg) remain available. In production these two series are produced by a job that reads a versioned price table. The public genaiotel lab follows a variant: it computes the cost of each span in the Collector (OTTL) from a price table, then a vmalert rule derives the cost series.

Grafana dashboards: one board per layer, correlated by trace_id, with as main panels p95 latency, error ratio by error.type, the distribution of finish_reasons, input tokens per request over time and cost per request and per feature. Design rule: every panel answers a decision, otherwise it is removed; a panel that represents an SLO has its alert (see the guide, 9.5).


Revised on 2 October 2026: prices removed, metric names fixed (_seconds suffix added by the exporter), TTFT rule added, content filter moved to the traces pipeline, pointer to the public genaiotel lab, mention of the acquisition of Langfuse by ClickHouse. Token metrics are evolving in the semantic-conventions-genai repository (see Part I).

Revised on 4 October 2026: the dashboard design rule points to guide 9.5, the corresponding pitfall having left Part VIII.