Skip to content

TechnicalPractitioner

Operations: retention, security, migration and FAQ

For: engineers and SREs · architectsPrerequisites: Have read lessons 1 to 5 of the course.

Module 12: optimization, retention and tuning

Section titled “Module 12: optimization, retention and tuning”

Not all LLM metrics deserve the same retention. The table gives recommended durations; the last column explains how to implement them.

Metric typeRecommended retentionJustificationImplementation
USD cost metricsexample: 2 to 3 years, to be set according to your obligationsAccounting audit, contractual SLAs“Long” instance (-retentionPeriod=3y) or Enterprise filter
RAG quality scores6 to 12 monthsLong-term drift detection“Medium” instance (-retentionPeriod=12M) or Enterprise filter
Latency and throughput3 to 6 monthsPerformance trending“Short” instance (-retentionPeriod=6M) or Enterprise filter
GPU metrics1 to 3 monthsCapacity planning“Short” instance or Enterprise filter
OTel traces7 to 30 daysDebugging, no long-term auditTTL of the trace backend (Jaeger, Tempo), outside VictoriaMetrics

Example of open source routing to two instances with vmagent:

Fenêtre de terminal
vmagent-prod \
-promscrape.config=/etc/prometheus.yml \
-remoteWrite.url=http://vm-long:8428/api/v1/write \
-remoteWrite.urlRelabelConfig=/etc/vmagent/keep-cost.yml \
-remoteWrite.url=http://vm-short:8428/api/v1/write \
-remoteWrite.urlRelabelConfig=/etc/vmagent/drop-cost.yml
# /etc/vmagent/keep-cost.yml: only cost and token series go to vm-long
- action: keep
source_labels: [__name__]
regex: 'llm_cost_usd_total|llm_tokens_total|llm:cost_usd:.*'
# /etc/vmagent/drop-cost.yml: everything else goes to vm-short
- action: drop
source_labels: [__name__]
regex: 'llm_cost_usd_total|llm_tokens_total|llm:cost_usd:.*'

Grafana then needs one datasource per instance, or a vmselect/vmauth layer in front if you want a single entry point.

  • Never label by trace_id, request_id or session_id: cardinality becomes unbounded.
  • Restrict tenant_id to cost metrics (not to latency or token metrics).
  • Group similar models under a model_family label if more than 20 versions are active.
  • Use histogram buckets adapted to tokens: [10, 100, 500, 1000, 4096, 16384, 65536].
# Cardinality statistics through the API:
# http://victoriametrics:8428/api/v1/status/tsdb
# (also available in vmui, "Cardinality explorer")
# Top 10 metrics by number of series (expensive on large instances)
topk(10, count by (__name__) ({__name__!=""}))

12.3 VictoriaMetrics tuning flags for LLM production

Section titled “12.3 VictoriaMetrics tuning flags for LLM production”

Start with the defaults and change a flag only after you have seen a need (rejected queries, rejected batches, memory pressure).

Fenêtre de terminal
# -loggerFormat=json: structured logs, easier to handle in a log pipeline
victoria-metrics-prod \
-storageDataPath=/var/lib/vm-data \
-retentionPeriod=12M \
-httpListenAddr=:8428 \
-loggerFormat=json

The following flags are often presented as tuning, but their usual values are already the defaults (flag list of VictoriaMetrics v1.153.0):

FlagDefaultWhen to change it
-maxInsertRequestSize33,554,432 bytes (32 MiB)If remote write batches are rejected as too large
-search.maxQueryLen16,384 bytesIf very long MetricsQL queries are rejected
-search.maxSamplesPerQuery1,000,000,000If audit queries over long windows hit the limit
-memory.allowedPercent60If caches run short or the OS page cache gets too small
-loggerLevelINFOTo reduce or increase log volume

Module 13: security, air-gap and compliance

Section titled “Module 13: security, air-gap and compliance”
  • VictoriaMetrics only offers a single global basic authentication (-httpAuth.username, -httpAuth.password) and per-endpoint keys. For per-user access control, put vmauth (open source) or a reverse proxy in front.
  • vmauth supports per-user routing: read-only for Grafana, write-only for the OTel Collector, admin for operations.
  • In Kubernetes, a NetworkPolicy restricts access to VictoriaMetrics ports to authorized components.
# vmauth config
users:
- username: grafana
password: '%{GRAFANA_PASSWORD}'
url_map:
- src_paths:
- '/api/v1/query'
- '/api/v1/query_range'
- '/api/v1/series'
- '/api/v1/labels'
- '/api/v1/label/.+/values'
url_prefix: 'http://victoriametrics:8428'
- username: otel-collector
password: '%{OTEL_PASSWORD}'
url_map:
- src_paths:
- '/api/v1/write'
- '/opentelemetry/.+'
url_prefix: 'http://victoriametrics:8428'
Compliance pointRecommended implementation
Personal data in metricsNever put user query content in labels. Use opaque identifiers only.
Cost auditRetention aligned with your accounting and contractual obligations on llm_cost_usd_total. Monthly CSV export.
Encryption in transitTLS on all VictoriaMetrics endpoints in production (vmauth or reverse proxy).
Multi-tenant isolationvmauth for per-tenant access segmentation, Grafana teams and RBAC.
BackupIncremental vmbackup to S3-compatible storage (MinIO in air-gap).
Dashboard accessGrafana with SSO (LDAP, OIDC). Viewer-only for business teams.

VictoriaMetrics can replace Prometheus as a storage and query backend in most cases. The migration below takes five steps; it is designed to avoid any service interruption and any data loss, provided each step is checked.

  • List all current Prometheus sources (scrape targets, remote write).
  • Measure current cardinality: promtool tsdb analyze /var/lib/prometheus/data/.
  • List all Grafana dashboards pointing to Prometheus.
  • List all Prometheus rules (alerting and recording).
  • Identify queries that rely on PromQL behaviour MetricsQL handles differently (see 14.4).

14.2 Dual-write strategy (to avoid downtime)

Section titled “14.2 Dual-write strategy (to avoid downtime)”

For two to four weeks, write metrics to both Prometheus and VictoriaMetrics. This lets you compare both backends without risk.

# prometheus.yml: add remote_write to VictoriaMetrics
remote_write:
- url: http://victoriametrics:8428/api/v1/write
queue_config:
capacity: 10000
max_shards: 30
max_samples_per_send: 5000
batch_send_deadline: 5s

14.3 Historical data import from Prometheus

Section titled “14.3 Historical data import from Prometheus”

To keep history, use vmctl (its options take a double dash). --vm-concurrency defaults to 2; leave --vm-batch-size at its default (200,000 samples) unless you have a measured need:

Fenêtre de terminal
# Import from a Prometheus snapshot
vmctl prometheus \
--prom-snapshot=/var/lib/prometheus/snapshots/20260101T000000Z-abc \
--vm-addr=http://victoriametrics:8428 \
--vm-concurrency=8
# Check in vmui:
# http://victoriametrics:8428/vmui/?g0.expr=count(up)

Two options depending on the size of the dashboard fleet:

  • Simple option: change the URL of the existing Prometheus datasource. Most dashboards keep working, since MetricsQL accepts PromQL syntax.
  • Recommended option: add a new VictoriaMetrics datasource (native plugin), migrate dashboards one by one, then remove the Prometheus datasource when dual-write stops.

Before switching Prometheus off, measure the same indicators on both backends during dual-write, with the same retention and the same queries:

IndicatorHow to measure it
RAM usedResident memory of the processes (process_resident_memory_bytes) over a week representative of your workload
Disk for the same retentionSize of the data directory, extrapolated to the target retention
P99 query latencyResponse time of Grafana panels and of the heaviest rules
Active series/api/v1/status/tsdb on VictoriaMetrics, promtool tsdb analyze on Prometheus
Monthly infrastructure costMachines, disks and object storage of both options, on the same scope

As an illustration, migrating four federated Prometheus instances to one VictoriaMetrics cluster can be planned over about six weeks, four of which in dual-write.

Scenario 1: internal RAG support chatbot (mid-sized company)

Section titled “Scenario 1: internal RAG support chatbot (mid-sized company)”

Context: tier-1 support assistant based on Mistral 7B hosted internally, weekly indexing of about 12,000 wiki documents, about 600 requests per day.

Metrics stack:

  • VictoriaMetrics Single (1 VM, 4 vCPU, 8 GB RAM, 250 GB SSD);
  • OTel Collector as a sidecar of the LLM API;
  • Grafana OSS, one dashboard, one read-only user for management;
  • vmalert with four rules, Microsoft Teams webhook output.

Estimated cardinality (order of magnitude):

Assumptions: a single model and a single provider, two pipelines (support questions and document search), eight indexes, the labels defined in lesson 03.

MetricCardinalityCalculation
llm_request_duration_ms24 serieslabels model, provider, pipeline_id (no status label): 1 x 1 x 2 = 2 combinations; per combination, 10 buckets (9 bounds plus +Inf) and the _sum and _count series, i.e. 12; 2 x 12 = 24
rag_retrieval_score_avg8 seriesone per index_id, each index serving a single pipeline with a single strategy
llm_tokens_total2 series1 model x 1 provider x 2 types (input, output)
DCGM_FI_DEV_GPU_UTIL2 seriesone per GPU, 2 L4 GPUs
Subtotal of listed metrics36 series24 + 8 + 2 + 2

With the other metrics of lesson 03 and the infrastructure exporters (DCGM, system), the total stays in the range of a few hundred to a few thousand active series, far below Single mode limits.

Key points:

  • With a few hundred to a few thousand series and a few hundred requests a day, Single mode is more than enough.
  • Recording rules lighten the cost panels, whose aggregation windows grow over the month.
  • In this scenario, the most relevant alert is RAGIndexDrift, to watch after each reindexing.

Scenario 2: multi-tenant LLM platform (B2B SaaS, about 40 customers)

Section titled “Scenario 2: multi-tenant LLM platform (B2B SaaS, about 40 customers)”

Context: conversational agent platform offered to about 40 companies, mixing OpenAI, Anthropic and an internal Llama 3 70B model on 8 H100 GPUs.

Metrics stack:

  • VictoriaMetrics Cluster (3 vmstorage, 2 vminsert, 2 vmselect);
  • one vmagent per region (EU-West, US-East);
  • Grafana with per-tenant teams and RBAC;
  • vmalert: 14 rules in total, not 14 per customer. Five of them (budget and quality, the metrics that carry tenant_id in this scenario) group their expression by tenant_id, like the budget rule in lesson 5: each is written once and fires a separate alert for each affected tenant.
RetentionMetricsJustification
3 yearsllm_cost_usd_total, llm_tokens_totalBilling, accounting audit
12 monthsrag_retrieval_score_avg, hallucination_rateLong-term quality drift detection
3 monthsDCGM_*, vllm:*, raw latencyShort-term capacity planning
7 days*_debug, *_tempIncident investigation

This tiered retention requires several storage groups or the Enterprise -retentionFilter (see 12.1).

Scenario 3: batch evaluation pipeline (R&D team)

Section titled “Scenario 3: batch evaluation pipeline (R&D team)”

Context: an ML team evaluates 12 prompt variants per week on 5 candidate models, 50,000 prompts per run, with LLM traces exported to a dedicated tracing and evaluation tool.

Architecture:

  • VictoriaMetrics Single in push mode (vmagent receives batches via remote write);
  • no Grafana: dashboards generated in Python (Plotly) after each run;
  • run metadata (run_id, prompt_template_id, model_version) carried by an info-style metric (see below);
  • short retention (90 days), daily export to a data warehouse for long-term analysis.

Q1. vmalert does not fire even with an obvious metric. Check three things. (1) Is the rule loaded? Call http://vmalert:8880/-/reload, then http://vmalert:8880/api/v1/rules. (2) Does the query return any series? Test it in vmui. (3) Is for: too long? An alert with for: 30m waits 30 minutes.

Q2. Cardinality explodes after adding a new service. Call http://victoriametrics:8428/api/v1/status/tsdb to identify the top labels. Often a trace_id or request_id label has leaked. Quick fix: a labeldrop rule in vmagent (-remoteWrite.relabelConfig).

Q3. My recording rules do not update.

Force a reload with curl http://vmalert:8880/-/reload (GET request), then check the logs for rule loading errors.

Q4. Difference between vmagent and vmauth? vmagent scrapes targets and sends data with remote write (the collection side of Prometheus). vmauth is a reverse proxy with authentication and per-user or per-tenant routing (useful for multi-tenancy and air-gap).

Q5. How do I know when to switch to Cluster mode? The vendor documentation recommends Single mode below about one million ingested data points per second and advises to think twice before moving to Cluster. So watch the ingestion rate first, then host saturation despite vertical scaling: RAM, CPU, query latency, disk. Alert thresholds on these resources are yours to set (as a rough indication, RAM persistently above 70% or P99 queries of several seconds); they are not vendor thresholds. A need for native high availability or multi-tenancy is the other reason to move to Cluster.

Q6. My histograms compute P99 incorrectly. Check the buckets: if the real P99 exceeds the last finite bucket, the result is clamped to that bucket. For LLMs, plan buckets up to 30 seconds: [50, 100, 250, 500, 1000, 2500, 5000, 10000, 30000] in milliseconds (or [.05, .1, .25, .5, 1, 2.5, 5, 10, 30] if your histogram is in seconds).

Q7. I lose samples in remote write from Prometheus. Tune queue_config in prometheus.yml: capacity, max_shards, max_samples_per_send. Watch prometheus_remote_storage_samples_dropped_total (or prometheus_remote_storage_samples_failed_total, depending on the version); if it is non-zero, raise max_shards.

Q8. How do I export data to a third-party platform? VictoriaMetrics exposes /api/v1/export (JSON lines). For continuous export, configure vmagent with two -remoteWrite.url, one to VictoriaMetrics and one to the third party. Watch the cost: every sample is billed by the third party as well.

Q9. Labels with non-ASCII characters cause problems. VictoriaMetrics supports UTF-8, but some tools render or escape these values poorly. Prefer ASCII values for business labels (for example tenant_id as a slug).

Q10. My dashboard takes 10 seconds to load after adding a variable. Each Grafana variable is a query run on refresh. Solutions: (1) refresh the variable only on dashboard load, (2) pre-aggregate with a recording rule, (3) limit the options (regex, sort, limit).

Q11. A GPU disappears from DCGM exporter metrics. DCGM exporter relies on NVML, which can lose the GPU list if the driver crashes. Check nvidia-smi on the node, then restart dcgm-exporter and nvidia-persistenced. To detect it, alert on absent(DCGM_FI_DEV_GPU_UTIL{node='X'}) (adapt the label to your relabelling).

Q12. How do I version my vmalert rules?

Store YAML files in Git, deploy with ArgoCD or Flux; vmalert reads them from a volume. Validate each pull request with the commands above.

Q13. rate() or increase() for costs? rate(llm_cost_usd_total[5m]) returns USD per second. increase(llm_cost_usd_total[1h]) returns the total increase over one hour, in USD. For a displayed hourly cost, prefer increase; for a projection, prefer rate.

Q14. My alerts send duplicates to Slack. Check group_by, group_wait and group_interval in Alertmanager. For LLMs, group_by: ['alertname', 'model'] groups alerts per model, and repeat_interval: 4h limits repetition.

Q15. Point-in-time backup: which strategy? vmbackup to S3-compatible storage (MinIO in air-gap), daily incremental snapshots, 30-day retention of backups, and a mandatory monthly restore test. Since backups are incremental, their storage cost stays a fraction of the live data size.