Technical
Why VictoriaMetrics, architecture and deployment mode
For: engineers and SREs · architectsPrerequisites: Prometheus basics (metric types, PromQL).
Module 1: why VictoriaMetrics for LLM workloads
Section titled “Module 1: why VictoriaMetrics for LLM workloads”1.1 The specific challenge of LLM metrics
Section titled “1.1 The specific challenge of LLM metrics”LLM workloads in production generate an atypical metrics profile that pushes classic monitoring backends to their limits.
| LLM characteristic | Impact on the metrics backend |
|---|---|
| High cardinality | model x provider x tenant x pipeline quickly yields thousands of time series, and every histogram bucket multiplies that number. |
| Token histograms | Wide distributions (10 to 100k tokens per request), with very spread-out buckets. |
| USD cost metrics | Financial metrics requiring float precision and aggregations by accounting period. |
| Long retention | Cost audits and contractual SLAs: one to three years of retention may be required. |
| Ingestion bursts | A completion batch can multiply the normal write rate within seconds; the size of the burst depends on the batch size and the scrape interval. |
| RAG quality metrics | Floats, distributions, scores: write-intensive and frequently read for alerting. |
1.2 VictoriaMetrics compared with Prometheus, Thanos and Mimir
Section titled “1.2 VictoriaMetrics compared with Prometheus, Thanos and Mimir”Prometheus on its own stores data on the local disk of one instance. For long retention, high availability or multi-tenancy, it is usually paired with remote storage: Thanos, Grafana Mimir (derived from Cortex) or VictoriaMetrics. The useful comparison therefore sets these options against each other, on verifiable criteria.
| Criterion | Prometheus alone | Prometheus + Thanos | Grafana Mimir | VictoriaMetrics |
|---|---|---|---|---|
| Deployment | Single binary | Prometheus plus several Thanos components (sidecar or receive, querier, store gateway, compactor) | Microservices, or monolithic mode in one binary | Single binary (Single), or three components (vminsert, vmselect, vmstorage) in Cluster |
| Storage and long retention | Local disk, not replicated (default retention: 15 days) | Object storage (S3, GCS, Azure…); downsampling by the compactor | Object storage | Local or block disk, retention per instance (-retentionPeriod); backups to object storage with vmbackup; downsampling restricted to the Enterprise edition |
| High availability | Two identical instances, no native deduplication | Deduplication at query time | Replication at ingestion | Two instances fed by vmagent (Single); replication factor in Cluster |
| Multi-tenancy | No | Yes, with Receive | Native | Native in Cluster (accountID) |
| Query language | PromQL | PromQL | PromQL | MetricsQL, largely PromQL-compatible, with extensions (median_over_time, rollup_candlestick, outliers_mad…) |
| OTLP ingestion | OTLP receiver (experimental since 2.47, built into Prometheus 3.x) | Through Prometheus or a Collector | Native OTLP endpoint | OTLP endpoint on the HTTP port: /opentelemetry/v1/metrics on port 8428 (Single), /insert/<accountID>/opentelemetry/v1/metrics on vminsert (Cluster) |
| License | Apache 2.0 | Apache 2.0 | AGPL 3.0 | Apache 2.0 (Single and Cluster); Enterprise edition under a commercial license (retention filters, downsampling, vmanomaly, vmauth mTLS…) |
| Compatibility | Reference | Prometheus query API | Prometheus query and write APIs | Prometheus query and write APIs: most existing dashboards work unchanged |
Hosted offerings (Grafana Cloud, Amazon Managed Service for Prometheus, Google Cloud Managed Service for Prometheus, VictoriaMetrics’ cloud offering, among others) build on these engines or on compatible APIs; they shift the question to cost per series and data location.
What the published figures say. The Prometheus documentation states an average of 1 to 2 bytes per sample (prometheus.io, Storage, accessed 2 October 2026). According to the vendor, VictoriaMetrics uses “up to 7x less RAM” than Prometheus, Thanos or Cortex under high cardinality (millions of unique time series) and requires “up to 7x less storage space” than the same tools (VictoriaMetrics documentation, accessed 2 October 2026). Both figures come from a benchmark published by the vendor itself, comparing Prometheus 2.22.2 and VictoriaMetrics 1.47.0 on node_exporter metrics (vendor article). They are maxima obtained on one specific workload with old versions: measure on your own workload, with the versions you plan to run, before sizing anything.
Module 2: complete VM + OTel architecture for LLM
Section titled “Module 2: complete VM + OTel architecture for LLM”2.1 Overview
Section titled “2.1 Overview”The architecture covers the full metrics path from LLM sources to dashboards and alerts. It can be deployed in a fully air-gapped environment. In summary, it chains the following stages:
| Stage | Components | What happens |
|---|---|---|
| Sources | LLM application (OTel SDK), vLLM, Ollama, LiteLLM, DCGM exporter | Emit metrics and traces, or expose a /metrics endpoint |
| Collection | OTel Collector, vmagent | Receive OTLP, scrape Prometheus endpoints, enrich, derive metrics from traces |
| Storage | VictoriaMetrics (Single or Cluster) | Stores series, serves MetricsQL, hosts vmui |
| Rules | vmalert, optionally vmanomaly | Evaluates alerting and recording rules, writes results back |
| Consumption | Grafana, Alertmanager | Dashboards, alert routing |
Key architectural points:
- Instrumented sources emit to the OTel Collector over OTLP (gRPC port 4317 or HTTP port 4318).
- The Collector filters, enriches (Kubernetes labels), derives metrics from traces (spanmetrics connector) and exports.
- VictoriaMetrics receives data via Prometheus remote write. It also accepts OTLP directly over HTTP, in both Single and Cluster mode.
- vmalert evaluates MetricsQL rules and sends alerts to Alertmanager.
- Grafana uses VictoriaMetrics as a datasource (VictoriaMetrics plugin or standard Prometheus datasource).
- vmanomaly (Enterprise) adds anomaly detection on LLM metrics.
2.2 Components and roles
Section titled “2.2 Components and roles”| Component | Port(s) | Role for LLM observability |
|---|---|---|
| OTel Collector | 4317 (gRPC), 4318 (HTTP) | Normalizes and routes all metrics. The spanmetrics connector derives metrics from RAG traces. |
| VictoriaMetrics Single | 8428 | TSDB storage, MetricsQL API, remote write, direct OTLP, built-in vmui |
| vmagent | 8429 | Prometheus-style scraping. Suited to DCGM exporter, vLLM /metrics, LiteLLM |
| vmalert | 8880 | Evaluates MetricsQL alerting and recording rules, talks to Alertmanager |
| Alertmanager | 9093 | Alert routing (Slack, PagerDuty, email, webhook), deduplication and silencing |
| Grafana | 3000 | Dashboards, with the VictoriaMetrics plugin or the standard Prometheus datasource |
| vmanomaly (Enterprise) | see its documentation | Anomaly detection on LLM metrics using statistical and ML models. Optional. |
Module 3: Single node vs Cluster, choosing your deployment mode
Section titled “Module 3: Single node vs Cluster, choosing your deployment mode”Single mode is one binary that ingests, stores and queries. Cluster mode splits these roles into three components: vminsert (ingestion), vmstorage (storage) and vmselect (queries), each of which can be scaled and replicated independently.
3.1 Decision criteria for an LLM context
Section titled “3.1 Decision criteria for an LLM context”The vendor documentation recommends the single-node version for ingestion rates lower than about a million data points per second, and advises to “think twice” before choosing the cluster version, which is harder to configure and operate (cluster documentation, accessed 2 October 2026). The LLM workloads described in this guide usually stay far below that rate: the choice of Cluster then mostly comes down to high availability and multi-tenancy.
| Criterion | Single mode | Cluster mode |
|---|---|---|
| Ingestion rate | Recommended below about one million data points per second (vendor documentation) | Above that, or when one machine is no longer enough |
| High availability | Not native (two instances fed by vmagent, or external solutions) | Native: replicas per component, replication factor |
| Multi-tenancy | Not native (separate instances or vmauth) | Native (accountID in API paths) |
| Deployment | 1 container, 1 binary | Three component types to deploy |
| Retention | One global retention per instance | One global retention per vmstorage; per-tenant or per-series retention requires Enterprise (-retentionFilter) |
| Ops complexity | Low | Higher: several components to size and monitor |
| License | Open source (Apache 2.0) | Open source (Apache 2.0) |
When to move from Single to Cluster
Section titled “When to move from Single to Cluster”Consider Cluster when ingestion approaches a million data points per second, when one machine’s resources are no longer enough despite vertical scaling, or when an HA SLA or native multi-tenant isolation is required. Queries, the Prometheus-compatible API and Grafana dashboards stay the same.