Skip to content

TechnicalPractitioner

Why VictoriaMetrics, architecture and deployment mode

For: engineers and SREs · architectsPrerequisites: Prometheus basics (metric types, PromQL).

Module 1: why VictoriaMetrics for LLM workloads

Section titled “Module 1: why VictoriaMetrics for LLM workloads”

LLM workloads in production generate an atypical metrics profile that pushes classic monitoring backends to their limits.

LLM characteristicImpact on the metrics backend
High cardinalitymodel x provider x tenant x pipeline quickly yields thousands of time series, and every histogram bucket multiplies that number.
Token histogramsWide distributions (10 to 100k tokens per request), with very spread-out buckets.
USD cost metricsFinancial metrics requiring float precision and aggregations by accounting period.
Long retentionCost audits and contractual SLAs: one to three years of retention may be required.
Ingestion burstsA completion batch can multiply the normal write rate within seconds; the size of the burst depends on the batch size and the scrape interval.
RAG quality metricsFloats, distributions, scores: write-intensive and frequently read for alerting.

1.2 VictoriaMetrics compared with Prometheus, Thanos and Mimir

Section titled “1.2 VictoriaMetrics compared with Prometheus, Thanos and Mimir”

Prometheus on its own stores data on the local disk of one instance. For long retention, high availability or multi-tenancy, it is usually paired with remote storage: Thanos, Grafana Mimir (derived from Cortex) or VictoriaMetrics. The useful comparison therefore sets these options against each other, on verifiable criteria.

CriterionPrometheus alonePrometheus + ThanosGrafana MimirVictoriaMetrics
DeploymentSingle binaryPrometheus plus several Thanos components (sidecar or receive, querier, store gateway, compactor)Microservices, or monolithic mode in one binarySingle binary (Single), or three components (vminsert, vmselect, vmstorage) in Cluster
Storage and long retentionLocal disk, not replicated (default retention: 15 days)Object storage (S3, GCS, Azure…); downsampling by the compactorObject storageLocal or block disk, retention per instance (-retentionPeriod); backups to object storage with vmbackup; downsampling restricted to the Enterprise edition
High availabilityTwo identical instances, no native deduplicationDeduplication at query timeReplication at ingestionTwo instances fed by vmagent (Single); replication factor in Cluster
Multi-tenancyNoYes, with ReceiveNativeNative in Cluster (accountID)
Query languagePromQLPromQLPromQLMetricsQL, largely PromQL-compatible, with extensions (median_over_time, rollup_candlestick, outliers_mad…)
OTLP ingestionOTLP receiver (experimental since 2.47, built into Prometheus 3.x)Through Prometheus or a CollectorNative OTLP endpointOTLP endpoint on the HTTP port: /opentelemetry/v1/metrics on port 8428 (Single), /insert/<accountID>/opentelemetry/v1/metrics on vminsert (Cluster)
LicenseApache 2.0Apache 2.0AGPL 3.0Apache 2.0 (Single and Cluster); Enterprise edition under a commercial license (retention filters, downsampling, vmanomaly, vmauth mTLS…)
CompatibilityReferencePrometheus query APIPrometheus query and write APIsPrometheus query and write APIs: most existing dashboards work unchanged

Hosted offerings (Grafana Cloud, Amazon Managed Service for Prometheus, Google Cloud Managed Service for Prometheus, VictoriaMetrics’ cloud offering, among others) build on these engines or on compatible APIs; they shift the question to cost per series and data location.

What the published figures say. The Prometheus documentation states an average of 1 to 2 bytes per sample (prometheus.io, Storage, accessed 2 October 2026). According to the vendor, VictoriaMetrics uses “up to 7x less RAM” than Prometheus, Thanos or Cortex under high cardinality (millions of unique time series) and requires “up to 7x less storage space” than the same tools (VictoriaMetrics documentation, accessed 2 October 2026). Both figures come from a benchmark published by the vendor itself, comparing Prometheus 2.22.2 and VictoriaMetrics 1.47.0 on node_exporter metrics (vendor article). They are maxima obtained on one specific workload with old versions: measure on your own workload, with the versions you plan to run, before sizing anything.

Module 2: complete VM + OTel architecture for LLM

Section titled “Module 2: complete VM + OTel architecture for LLM”

The architecture covers the full metrics path from LLM sources to dashboards and alerts. It can be deployed in a fully air-gapped environment. In summary, it chains the following stages:

StageComponentsWhat happens
SourcesLLM application (OTel SDK), vLLM, Ollama, LiteLLM, DCGM exporterEmit metrics and traces, or expose a /metrics endpoint
CollectionOTel Collector, vmagentReceive OTLP, scrape Prometheus endpoints, enrich, derive metrics from traces
StorageVictoriaMetrics (Single or Cluster)Stores series, serves MetricsQL, hosts vmui
Rulesvmalert, optionally vmanomalyEvaluates alerting and recording rules, writes results back
ConsumptionGrafana, AlertmanagerDashboards, alert routing

Key architectural points:

  • Instrumented sources emit to the OTel Collector over OTLP (gRPC port 4317 or HTTP port 4318).
  • The Collector filters, enriches (Kubernetes labels), derives metrics from traces (spanmetrics connector) and exports.
  • VictoriaMetrics receives data via Prometheus remote write. It also accepts OTLP directly over HTTP, in both Single and Cluster mode.
  • vmalert evaluates MetricsQL rules and sends alerts to Alertmanager.
  • Grafana uses VictoriaMetrics as a datasource (VictoriaMetrics plugin or standard Prometheus datasource).
  • vmanomaly (Enterprise) adds anomaly detection on LLM metrics.
ComponentPort(s)Role for LLM observability
OTel Collector4317 (gRPC), 4318 (HTTP)Normalizes and routes all metrics. The spanmetrics connector derives metrics from RAG traces.
VictoriaMetrics Single8428TSDB storage, MetricsQL API, remote write, direct OTLP, built-in vmui
vmagent8429Prometheus-style scraping. Suited to DCGM exporter, vLLM /metrics, LiteLLM
vmalert8880Evaluates MetricsQL alerting and recording rules, talks to Alertmanager
Alertmanager9093Alert routing (Slack, PagerDuty, email, webhook), deduplication and silencing
Grafana3000Dashboards, with the VictoriaMetrics plugin or the standard Prometheus datasource
vmanomaly (Enterprise)see its documentationAnomaly detection on LLM metrics using statistical and ML models. Optional.

Module 3: Single node vs Cluster, choosing your deployment mode

Section titled “Module 3: Single node vs Cluster, choosing your deployment mode”

Single mode is one binary that ingests, stores and queries. Cluster mode splits these roles into three components: vminsert (ingestion), vmstorage (storage) and vmselect (queries), each of which can be scaled and replicated independently.

The vendor documentation recommends the single-node version for ingestion rates lower than about a million data points per second, and advises to “think twice” before choosing the cluster version, which is harder to configure and operate (cluster documentation, accessed 2 October 2026). The LLM workloads described in this guide usually stay far below that rate: the choice of Cluster then mostly comes down to high availability and multi-tenancy.

CriterionSingle modeCluster mode
Ingestion rateRecommended below about one million data points per second (vendor documentation)Above that, or when one machine is no longer enough
High availabilityNot native (two instances fed by vmagent, or external solutions)Native: replicas per component, replication factor
Multi-tenancyNot native (separate instances or vmauth)Native (accountID in API paths)
Deployment1 container, 1 binaryThree component types to deploy
RetentionOne global retention per instanceOne global retention per vmstorage; per-tenant or per-series retention requires Enterprise (-retentionFilter)
Ops complexityLowHigher: several components to size and monitor
LicenseOpen source (Apache 2.0)Open source (Apache 2.0)

Consider Cluster when ingestion approaches a million data points per second, when one machine’s resources are no longer enough despite vertical scaling, or when an HA SLA or native multi-tenant isolation is required. Queries, the Prometheus-compatible API and Grafana dashboards stay the same.