Technical
VictoriaMetrics as an LLM observability backend
For: engineers and SREs · architectsPrerequisites: Prometheus basics (metric types, PromQL), Docker Compose and some Python.
This course completes the “Go to production” step: it covers one component of the production stack (the metrics backend) and comes after the labs and the deployment playbooks.
Audience: SRE, platform and MLOps engineers, observability architects. Prerequisites: basic Prometheus knowledge (metric types, PromQL), Docker Compose, some Python. Goal: run VictoriaMetrics as a production metrics backend for LLM and RAG workloads, from first deployment to alerting and SLOs.
This guide covers a concrete setup: OpenTelemetry Collector and vmagent for collection, VictoriaMetrics for storage, MetricsQL for queries, Grafana for dashboards, vmalert and Alertmanager for alerts. The architecture can be deployed in an air-gapped environment.
This course takes up and corrects my guide “VictoriaMetrics as an LLM Observability Backend” (version 3). Callouts in the lessons flag the points corrected since that first version. The labs rely on the public VMLLM lab on the forge.
Learning objectives
Section titled “Learning objectives”By the end of this guide, you will be able to:
- deploy VictoriaMetrics in Single mode for an LLM development environment;
- choose between Single and Cluster mode based on your actual LLM workload;
- configure the OTel Collector to VictoriaMetrics pipeline for LLM metrics collection;
- instrument a Python/FastAPI application to emit the right LLM metrics;
- write MetricsQL queries for the main use cases (cost, quality, latency, GPU);
- build a Grafana LLM dashboard with four functional rows and real-time alerts;
- configure vmalert with production-grade alert rules for LLM degradations;
- optimize retention, storage and recording rules to minimize infrastructure costs.
Contents
Section titled “Contents”- Why VictoriaMetrics, architecture and deployment mode: the LLM metrics profile, comparison with Prometheus, Thanos and Mimir, the VM + OTel architecture, Single versus Cluster.
- Installation and configuration: Docker Compose stack, OTel Collector with the spanmetrics connector, vmagent scraping.
- Metrics and instrumentation: LLM metrics taxonomy, naming, cardinality, Python instrumentation.
- MetricsQL and dashboards: cost, quality, latency and GPU queries, Grafana provisioning and panels.
- Alerts, recording rules and SLOs: vmalert ruleset, Alertmanager routing, recording rules, error budget and burn rate.
- Operations: retention and tuning, security and air-gap, migration from Prometheus, illustrative scenarios, FAQ.
- Labs, checklist and cheat sheet: four practical labs, go-live checklist, MetricsQL cheat sheet, references.
- Glossary.
Metric names and GenAI conventions
Section titled “Metric names and GenAI conventions”This course names its application metrics llm_* and rag_* and emits them with the Prometheus client. These are not the names of the OpenTelemetry GenAI semantic conventions, whose status and history are kept in the article, section 2. The correspondence principle: the quantities are the same, the names and units differ. If your stack also receives gen_ai.* metrics from an OpenTelemetry instrumentation, keep a single source per quantity and do the translation in the Collector or in recording rules rather than in every dashboard. The course names stay unchanged: they match the lab code.
| Course metric | Equivalent in the GenAI conventions | Note |
|---|---|---|
llm_requests_total | number of observations of the gen_ai.client.operation.duration histogram | failure status is carried by error.type |
llm_request_duration_ms | gen_ai.client.operation.duration | the conventions measure in seconds |
llm_ttft_ms | gen_ai.client.operation.time_to_first_chunk (client side) or gen_ai.server.time_to_first_token (server side) | in seconds |
llm_tokens_total{type} | gen_ai.client.token.usage broken down by gen_ai.token.type; on the unreleased main branch of the GenAI repository, gen_ai.client.inference.usage.input_tokens and gen_ai.client.inference.usage.output_tokens | see the article, section 2.1 |
llm_errors_total | error.type attribute on the duration metrics | |
llm_cost_usd_total | none: cost is not part of the conventions | application metric, keep it |
rag_* | no metric; retrieval spans with gen_ai.data_source.id and gen_ai.retrieval.top_k | application metrics |
model, provider labels | gen_ai.request.model, gen_ai.provider.name attributes |
Going further
Section titled “Going further”- Understanding generative AI observability: concepts and definitions.
- LLM observability: the labs: hands-on, on a self-hosted stack.
- Observing an LLM system: deployment playbooks and the OpenTelemetry GenAI conventions.
- GenAI observability method: quality SLOs, quantified impact, compliance.
Revised on 4 October 2026: page moved to MDX with the progression box, positioning within the production stack, mapping between the llm_* metrics and the GenAI conventions, links to the guide, the labs, the article and the method.