Technical
Operations: retention, security, migration and FAQ
For: engineers and SREs · architectsPrerequisites: Have read lessons 1 to 5 of the course.
Module 12: optimization, retention and tuning
Section titled “Module 12: optimization, retention and tuning”12.1 Retention policy by metric type
Section titled “12.1 Retention policy by metric type”Not all LLM metrics deserve the same retention. The table gives recommended durations; the last column explains how to implement them.
| Metric type | Recommended retention | Justification | Implementation |
|---|---|---|---|
| USD cost metrics | example: 2 to 3 years, to be set according to your obligations | Accounting audit, contractual SLAs | “Long” instance (-retentionPeriod=3y) or Enterprise filter |
| RAG quality scores | 6 to 12 months | Long-term drift detection | “Medium” instance (-retentionPeriod=12M) or Enterprise filter |
| Latency and throughput | 3 to 6 months | Performance trending | “Short” instance (-retentionPeriod=6M) or Enterprise filter |
| GPU metrics | 1 to 3 months | Capacity planning | “Short” instance or Enterprise filter |
| OTel traces | 7 to 30 days | Debugging, no long-term audit | TTL of the trace backend (Jaeger, Tempo), outside VictoriaMetrics |
Example of open source routing to two instances with vmagent:
vmagent-prod \ -promscrape.config=/etc/prometheus.yml \ -remoteWrite.url=http://vm-long:8428/api/v1/write \ -remoteWrite.urlRelabelConfig=/etc/vmagent/keep-cost.yml \ -remoteWrite.url=http://vm-short:8428/api/v1/write \ -remoteWrite.urlRelabelConfig=/etc/vmagent/drop-cost.yml# /etc/vmagent/keep-cost.yml: only cost and token series go to vm-long- action: keep source_labels: [__name__] regex: 'llm_cost_usd_total|llm_tokens_total|llm:cost_usd:.*'# /etc/vmagent/drop-cost.yml: everything else goes to vm-short- action: drop source_labels: [__name__] regex: 'llm_cost_usd_total|llm_tokens_total|llm:cost_usd:.*'Grafana then needs one datasource per instance, or a vmselect/vmauth layer in front if you want a single entry point.
12.2 Cardinality optimization
Section titled “12.2 Cardinality optimization”- Never label by
trace_id,request_idorsession_id: cardinality becomes unbounded. - Restrict
tenant_idto cost metrics (not to latency or token metrics). - Group similar models under a
model_familylabel if more than 20 versions are active. - Use histogram buckets adapted to tokens:
[10, 100, 500, 1000, 4096, 16384, 65536].
# Cardinality statistics through the API:# http://victoriametrics:8428/api/v1/status/tsdb# (also available in vmui, "Cardinality explorer")
# Top 10 metrics by number of series (expensive on large instances)topk(10, count by (__name__) ({__name__!=""}))12.3 VictoriaMetrics tuning flags for LLM production
Section titled “12.3 VictoriaMetrics tuning flags for LLM production”Start with the defaults and change a flag only after you have seen a need (rejected queries, rejected batches, memory pressure).
# -loggerFormat=json: structured logs, easier to handle in a log pipelinevictoria-metrics-prod \ -storageDataPath=/var/lib/vm-data \ -retentionPeriod=12M \ -httpListenAddr=:8428 \ -loggerFormat=jsonThe following flags are often presented as tuning, but their usual values are already the defaults (flag list of VictoriaMetrics v1.153.0):
| Flag | Default | When to change it |
|---|---|---|
-maxInsertRequestSize | 33,554,432 bytes (32 MiB) | If remote write batches are rejected as too large |
-search.maxQueryLen | 16,384 bytes | If very long MetricsQL queries are rejected |
-search.maxSamplesPerQuery | 1,000,000,000 | If audit queries over long windows hit the limit |
-memory.allowedPercent | 60 | If caches run short or the OS page cache gets too small |
-loggerLevel | INFO | To reduce or increase log volume |
Module 13: security, air-gap and compliance
Section titled “Module 13: security, air-gap and compliance”13.1 Authentication and access control
Section titled “13.1 Authentication and access control”- VictoriaMetrics only offers a single global basic authentication (
-httpAuth.username,-httpAuth.password) and per-endpoint keys. For per-user access control, put vmauth (open source) or a reverse proxy in front. - vmauth supports per-user routing: read-only for Grafana, write-only for the OTel Collector, admin for operations.
- In Kubernetes, a NetworkPolicy restricts access to VictoriaMetrics ports to authorized components.
# vmauth configusers: - username: grafana password: '%{GRAFANA_PASSWORD}' url_map: - src_paths: - '/api/v1/query' - '/api/v1/query_range' - '/api/v1/series' - '/api/v1/labels' - '/api/v1/label/.+/values' url_prefix: 'http://victoriametrics:8428'
- username: otel-collector password: '%{OTEL_PASSWORD}' url_map: - src_paths: - '/api/v1/write' - '/opentelemetry/.+' url_prefix: 'http://victoriametrics:8428'13.2 Compliance checklist for LLM metrics
Section titled “13.2 Compliance checklist for LLM metrics”| Compliance point | Recommended implementation |
|---|---|
| Personal data in metrics | Never put user query content in labels. Use opaque identifiers only. |
| Cost audit | Retention aligned with your accounting and contractual obligations on llm_cost_usd_total. Monthly CSV export. |
| Encryption in transit | TLS on all VictoriaMetrics endpoints in production (vmauth or reverse proxy). |
| Multi-tenant isolation | vmauth for per-tenant access segmentation, Grafana teams and RBAC. |
| Backup | Incremental vmbackup to S3-compatible storage (MinIO in air-gap). |
| Dashboard access | Grafana with SSO (LDAP, OIDC). Viewer-only for business teams. |
Module 14: migrating from Prometheus
Section titled “Module 14: migrating from Prometheus”VictoriaMetrics can replace Prometheus as a storage and query backend in most cases. The migration below takes five steps; it is designed to avoid any service interruption and any data loss, provided each step is checked.
14.1 Pre-migration inventory
Section titled “14.1 Pre-migration inventory”- List all current Prometheus sources (scrape targets, remote write).
- Measure current cardinality:
promtool tsdb analyze /var/lib/prometheus/data/. - List all Grafana dashboards pointing to Prometheus.
- List all Prometheus rules (alerting and recording).
- Identify queries that rely on PromQL behaviour MetricsQL handles differently (see 14.4).
14.2 Dual-write strategy (to avoid downtime)
Section titled “14.2 Dual-write strategy (to avoid downtime)”For two to four weeks, write metrics to both Prometheus and VictoriaMetrics. This lets you compare both backends without risk.
# prometheus.yml: add remote_write to VictoriaMetricsremote_write: - url: http://victoriametrics:8428/api/v1/write queue_config: capacity: 10000 max_shards: 30 max_samples_per_send: 5000 batch_send_deadline: 5s14.3 Historical data import from Prometheus
Section titled “14.3 Historical data import from Prometheus”To keep history, use vmctl (its options take a double dash). --vm-concurrency defaults to 2; leave --vm-batch-size at its default (200,000 samples) unless you have a measured need:
# Import from a Prometheus snapshotvmctl prometheus \ --prom-snapshot=/var/lib/prometheus/snapshots/20260101T000000Z-abc \ --vm-addr=http://victoriametrics:8428 \ --vm-concurrency=8
# Check in vmui:# http://victoriametrics:8428/vmui/?g0.expr=count(up)14.4 Grafana dashboard migration
Section titled “14.4 Grafana dashboard migration”Two options depending on the size of the dashboard fleet:
- Simple option: change the URL of the existing Prometheus datasource. Most dashboards keep working, since MetricsQL accepts PromQL syntax.
- Recommended option: add a new VictoriaMetrics datasource (native plugin), migrate dashboards one by one, then remove the Prometheus datasource when dual-write stops.
14.5 Final cutover and gains to measure
Section titled “14.5 Final cutover and gains to measure”Before switching Prometheus off, measure the same indicators on both backends during dual-write, with the same retention and the same queries:
| Indicator | How to measure it |
|---|---|
| RAM used | Resident memory of the processes (process_resident_memory_bytes) over a week representative of your workload |
| Disk for the same retention | Size of the data directory, extrapolated to the target retention |
| P99 query latency | Response time of Grafana panels and of the heaviest rules |
| Active series | /api/v1/status/tsdb on VictoriaMetrics, promtool tsdb analyze on Prometheus |
| Monthly infrastructure cost | Machines, disks and object storage of both options, on the same scope |
As an illustration, migrating four federated Prometheus instances to one VictoriaMetrics cluster can be planned over about six weeks, four of which in dual-write.
Module 15: illustrative scenarios
Section titled “Module 15: illustrative scenarios”Scenario 1: internal RAG support chatbot (mid-sized company)
Section titled “Scenario 1: internal RAG support chatbot (mid-sized company)”Context: tier-1 support assistant based on Mistral 7B hosted internally, weekly indexing of about 12,000 wiki documents, about 600 requests per day.
Metrics stack:
- VictoriaMetrics Single (1 VM, 4 vCPU, 8 GB RAM, 250 GB SSD);
- OTel Collector as a sidecar of the LLM API;
- Grafana OSS, one dashboard, one read-only user for management;
- vmalert with four rules, Microsoft Teams webhook output.
Estimated cardinality (order of magnitude):
Assumptions: a single model and a single provider, two pipelines (support questions and document search), eight indexes, the labels defined in lesson 03.
| Metric | Cardinality | Calculation |
|---|---|---|
llm_request_duration_ms | 24 series | labels model, provider, pipeline_id (no status label): 1 x 1 x 2 = 2 combinations; per combination, 10 buckets (9 bounds plus +Inf) and the _sum and _count series, i.e. 12; 2 x 12 = 24 |
rag_retrieval_score_avg | 8 series | one per index_id, each index serving a single pipeline with a single strategy |
llm_tokens_total | 2 series | 1 model x 1 provider x 2 types (input, output) |
DCGM_FI_DEV_GPU_UTIL | 2 series | one per GPU, 2 L4 GPUs |
| Subtotal of listed metrics | 36 series | 24 + 8 + 2 + 2 |
With the other metrics of lesson 03 and the infrastructure exporters (DCGM, system), the total stays in the range of a few hundred to a few thousand active series, far below Single mode limits.
Key points:
- With a few hundred to a few thousand series and a few hundred requests a day, Single mode is more than enough.
- Recording rules lighten the cost panels, whose aggregation windows grow over the month.
- In this scenario, the most relevant alert is
RAGIndexDrift, to watch after each reindexing.
Scenario 2: multi-tenant LLM platform (B2B SaaS, about 40 customers)
Section titled “Scenario 2: multi-tenant LLM platform (B2B SaaS, about 40 customers)”Context: conversational agent platform offered to about 40 companies, mixing OpenAI, Anthropic and an internal Llama 3 70B model on 8 H100 GPUs.
Metrics stack:
- VictoriaMetrics Cluster (3 vmstorage, 2 vminsert, 2 vmselect);
- one vmagent per region (EU-West, US-East);
- Grafana with per-tenant teams and RBAC;
- vmalert: 14 rules in total, not 14 per customer. Five of them (budget and quality, the metrics that carry
tenant_idin this scenario) group their expression bytenant_id, like the budget rule in lesson 5: each is written once and fires a separate alert for each affected tenant.
| Retention | Metrics | Justification |
|---|---|---|
| 3 years | llm_cost_usd_total, llm_tokens_total | Billing, accounting audit |
| 12 months | rag_retrieval_score_avg, hallucination_rate | Long-term quality drift detection |
| 3 months | DCGM_*, vllm:*, raw latency | Short-term capacity planning |
| 7 days | *_debug, *_temp | Incident investigation |
This tiered retention requires several storage groups or the Enterprise -retentionFilter (see 12.1).
Scenario 3: batch evaluation pipeline (R&D team)
Section titled “Scenario 3: batch evaluation pipeline (R&D team)”Context: an ML team evaluates 12 prompt variants per week on 5 candidate models, 50,000 prompts per run, with LLM traces exported to a dedicated tracing and evaluation tool.
Architecture:
- VictoriaMetrics Single in push mode (vmagent receives batches via remote write);
- no Grafana: dashboards generated in Python (Plotly) after each run;
- run metadata (
run_id,prompt_template_id,model_version) carried by an info-style metric (see below); - short retention (90 days), daily export to a data warehouse for long-term analysis.
Module 16: FAQ and troubleshooting
Section titled “Module 16: FAQ and troubleshooting”Q1. vmalert does not fire even with an obvious metric. Check three things. (1) Is the rule loaded? Call http://vmalert:8880/-/reload, then http://vmalert:8880/api/v1/rules. (2) Does the query return any series? Test it in vmui. (3) Is for: too long? An alert with for: 30m waits 30 minutes.
Q2. Cardinality explodes after adding a new service. Call http://victoriametrics:8428/api/v1/status/tsdb to identify the top labels. Often a trace_id or request_id label has leaked. Quick fix: a labeldrop rule in vmagent (-remoteWrite.relabelConfig).
Q3. My recording rules do not update.
Force a reload with curl http://vmalert:8880/-/reload (GET request), then check the logs for rule loading errors.
Q4. Difference between vmagent and vmauth? vmagent scrapes targets and sends data with remote write (the collection side of Prometheus). vmauth is a reverse proxy with authentication and per-user or per-tenant routing (useful for multi-tenancy and air-gap).
Q5. How do I know when to switch to Cluster mode? The vendor documentation recommends Single mode below about one million ingested data points per second and advises to think twice before moving to Cluster. So watch the ingestion rate first, then host saturation despite vertical scaling: RAM, CPU, query latency, disk. Alert thresholds on these resources are yours to set (as a rough indication, RAM persistently above 70% or P99 queries of several seconds); they are not vendor thresholds. A need for native high availability or multi-tenancy is the other reason to move to Cluster.
Q6. My histograms compute P99 incorrectly. Check the buckets: if the real P99 exceeds the last finite bucket, the result is clamped to that bucket. For LLMs, plan buckets up to 30 seconds: [50, 100, 250, 500, 1000, 2500, 5000, 10000, 30000] in milliseconds (or [.05, .1, .25, .5, 1, 2.5, 5, 10, 30] if your histogram is in seconds).
Q7. I lose samples in remote write from Prometheus. Tune queue_config in prometheus.yml: capacity, max_shards, max_samples_per_send. Watch prometheus_remote_storage_samples_dropped_total (or prometheus_remote_storage_samples_failed_total, depending on the version); if it is non-zero, raise max_shards.
Q8. How do I export data to a third-party platform? VictoriaMetrics exposes /api/v1/export (JSON lines). For continuous export, configure vmagent with two -remoteWrite.url, one to VictoriaMetrics and one to the third party. Watch the cost: every sample is billed by the third party as well.
Q9. Labels with non-ASCII characters cause problems. VictoriaMetrics supports UTF-8, but some tools render or escape these values poorly. Prefer ASCII values for business labels (for example tenant_id as a slug).
Q10. My dashboard takes 10 seconds to load after adding a variable. Each Grafana variable is a query run on refresh. Solutions: (1) refresh the variable only on dashboard load, (2) pre-aggregate with a recording rule, (3) limit the options (regex, sort, limit).
Q11. A GPU disappears from DCGM exporter metrics. DCGM exporter relies on NVML, which can lose the GPU list if the driver crashes. Check nvidia-smi on the node, then restart dcgm-exporter and nvidia-persistenced. To detect it, alert on absent(DCGM_FI_DEV_GPU_UTIL{node='X'}) (adapt the label to your relabelling).
Q12. How do I version my vmalert rules?
Store YAML files in Git, deploy with ArgoCD or Flux; vmalert reads them from a volume. Validate each pull request with the commands above.
Q13. rate() or increase() for costs? rate(llm_cost_usd_total[5m]) returns USD per second. increase(llm_cost_usd_total[1h]) returns the total increase over one hour, in USD. For a displayed hourly cost, prefer increase; for a projection, prefer rate.
Q14. My alerts send duplicates to Slack. Check group_by, group_wait and group_interval in Alertmanager. For LLMs, group_by: ['alertname', 'model'] groups alerts per model, and repeat_interval: 4h limits repetition.
Q15. Point-in-time backup: which strategy? vmbackup to S3-compatible storage (MinIO in air-gap), daily incremental snapshots, 30-day retention of backups, and a mandatory monthly restore test. Since backups are incremental, their storage cost stays a fraction of the live data size.