Skip to content

TechnicalPractitioner

Installation and configuration step by step

For: engineers and SREsPrerequisites: Have read lesson 1 of the course and know Docker Compose.

A development or lightweight production stack: VictoriaMetrics Single, OTel Collector, vmagent, vmalert, Alertmanager and Grafana.

docker-compose.yml
services:
victoriametrics:
image: victoriametrics/victoria-metrics:v1.153.0
command:
- '-storageDataPath=/var/lib/victoria-metrics-data'
- '-retentionPeriod=12M' # 12 months (upper-case M suffix)
- '-httpListenAddr=:8428'
- '-opentelemetry.usePrometheusNaming' # Prometheus names for direct OTLP (see 4.2)
ports: ['8428:8428']
volumes:
- vm-data:/var/lib/victoria-metrics-data
otel-collector:
image: otel/opentelemetry-collector-contrib:0.161.0
command: ['--config=/etc/otel-collector-config.yaml']
volumes:
- ./otel-collector-config.yaml:/etc/otel-collector-config.yaml
ports:
- '4317:4317' # OTLP gRPC
- '4318:4318' # OTLP HTTP
depends_on: [victoriametrics]
vmagent:
image: victoriametrics/vmagent:v1.153.0
command:
- '-promscrape.config=/etc/prometheus.yml'
- '-remoteWrite.url=http://victoriametrics:8428/api/v1/write'
volumes:
- ./prometheus.yml:/etc/prometheus.yml
vmalert:
image: victoriametrics/vmalert:v1.153.0
command:
- '-datasource.url=http://victoriametrics:8428'
- '-remoteWrite.url=http://victoriametrics:8428/api/v1/write'
- '-notifier.url=http://alertmanager:9093'
- '-rule=/etc/alerts/*.yml'
volumes:
- ./alerts:/etc/alerts
ports: ['8880:8880']
alertmanager:
image: prom/alertmanager:v0.34.1
volumes: ['./alertmanager.yml:/etc/alertmanager/alertmanager.yml']
ports: ['9093:9093']
grafana:
image: grafana/grafana:12.4.12
environment:
- GF_SECURITY_ADMIN_PASSWORD=admin # development only, change it
- GF_PLUGINS_PREINSTALL=victoriametrics-metrics-datasource
ports: ['3000:3000']
volumes:
vm-data: {}

Notes on this file:

  • -retentionPeriod without a suffix is expressed in months (of 31 days). The documentation lists the suffixes s, h, d, w, M (month, upper case) and y, for example 90d, 12M or 2y. The suffix-less form is accepted for historical reasons, but the VictoriaMetrics source code marks it as deprecated: prefer M. Do not write a lower-case m for months: it means minutes and VictoriaMetrics rejects it for this flag. The default is 1M (one month) and the minimum is 1d (or 24h).
  • -remoteWrite.url of vmalert is used to store the results of recording rules and alert states in VictoriaMetrics.
  • Port 8880 of vmalert is published so that its web UI can be used in the labs.
  • -opentelemetry.usePrometheusNaming only matters when metrics reach VictoriaMetrics directly over OTLP (see 4.2). Without it, OTLP names are stored as is.
  • GF_PLUGINS_PREINSTALL replaces GF_INSTALL_PLUGINS, deprecated since Grafana 12.1.
  • The Grafana admin password admin is acceptable only on a workstation.

4.2 OTel Collector configuration for LLM metrics

Section titled “4.2 OTel Collector configuration for LLM metrics”
otel-collector-config.yaml
receivers:
otlp:
protocols:
grpc: { endpoint: 0.0.0.0:4317 }
http: { endpoint: 0.0.0.0:4318 }
prometheus:
config:
scrape_configs:
- job_name: vllm
scrape_interval: 15s
static_configs: [{ targets: ['vllm-service:8000'] }]
processors:
batch: { timeout: 5s, send_batch_size: 512 }
k8sattributes:
extract:
metadata: [k8s.pod.name, k8s.namespace.name, k8s.deployment.name]
connectors:
spanmetrics:
histogram:
explicit:
buckets: [50ms, 100ms, 200ms, 500ms, 1s, 2s, 5s, 10s]
dimensions:
- name: gen_ai.request.model
- name: gen_ai.provider.name
- name: rag.index.id
- name: rag.pipeline.id
- name: rag.tenant.id
exporters:
prometheusremotewrite:
endpoint: http://victoriametrics:8428/api/v1/write
timeout: 30s
retry_on_failure: { enabled: true, initial_interval: 5s }
service:
pipelines:
traces:
receivers: [otlp]
processors: [k8sattributes, batch]
exporters: [spanmetrics] # add your trace backend exporter here
metrics:
receivers: [otlp, prometheus, spanmetrics]
processors: [k8sattributes, batch]
exporters: [prometheusremotewrite]

Points to check:

  • k8sattributes and RBAC. In a Kubernetes cluster, the k8sattributes processor queries the API server. Its service account needs a ClusterRole allowing get, list and watch on pods and namespaces, and on replicasets to resolve k8s.deployment.name. Outside Kubernetes (Docker Compose on a workstation), remove it from the pipelines.
  • Trace backend. The traces pipeline above only feeds spanmetrics. In practice, also add an exporter to your trace backend (Jaeger, Tempo or another OTLP endpoint).
  • Names of derived metrics. The names produced by the connector (calls counter, duration histogram) depend on the Collector version and on the connector’s namespace option. Check them in vmui before writing dashboards.
  • Alternative exporter. Instead of prometheusremotewrite, you can send OTLP directly to VictoriaMetrics with the otlphttp exporter pointing to http://victoriametrics:8428/opentelemetry. In that case, start VictoriaMetrics (or the vmagent receiving the OTLP data) with -opentelemetry.usePrometheusNaming, as in the Compose file above. Without this flag, VictoriaMetrics stores OTLP points without any transformation: no _total suffix and no unit, dots in names and labels, and every *_total and *_bucket query in this guide stops working (VictoriaMetrics OpenTelemetry documentation).
  • Dimensions. gen_ai.request.model and gen_ai.provider.name come from the OpenTelemetry GenAI semantic conventions, moved to the semantic-conventions-genai repository since v1.42 (June 2026); gen_ai.provider.name replaces the former gen_ai.system. rag.* attributes are application-specific conventions of this guide.

4.3 vmagent: scraping DCGM and LLM endpoints

Section titled “4.3 vmagent: scraping DCGM and LLM endpoints”
# prometheus.yml (used by vmagent)
global:
scrape_interval: 15s
external_labels:
cluster: 'prod-llm'
environment: 'production'
scrape_configs:
# DCGM exporter: NVIDIA GPU metrics
- job_name: 'dcgm-exporter'
scrape_interval: 5s
kubernetes_sd_configs: [{ role: pod }]
relabel_configs:
- source_labels: [__meta_kubernetes_pod_label_app]
regex: dcgm-exporter
action: keep
# vLLM inference servers
- job_name: 'vllm'
static_configs:
- targets: ['vllm-llama:8000', 'vllm-mistral:8001']
# LiteLLM proxy
- job_name: 'litellm'
static_configs: [{ targets: ['litellm-proxy:4000'] }]

The kubernetes_sd_configs block only works when vmagent runs in Kubernetes with the right permissions. In Docker Compose, replace it with a static_configs entry pointing to the DCGM exporter (port 9400 by default).

Revised on 2 October 2026: images pinned to current releases, GF_PLUGINS_PREINSTALL instead of GF_INSTALL_PLUGINS, gen_ai.provider.name instead of gen_ai.system, -opentelemetry.usePrometheusNaming flag for direct OTLP ingestion, retention suffixes aligned with the documentation.