Technical
Labs, go-live checklist and MetricsQL cheat sheet
For: engineers and SREsPrerequisites: Have read lessons 1 to 5 of the course and have Docker with Compose v2.
Module 17: practical labs
Section titled “Module 17: practical labs”Lab 1: deploy the stack and ingest first metrics
Section titled “Lab 1: deploy the stack and ingest first metrics”Goal: deploy the complete Docker Compose stack (VictoriaMetrics, OTel Collector, vmagent, vmalert, Alertmanager, Grafana, simulator), check the metrics in vmui and Grafana. Estimated time: 30 minutes. Level: beginner. Prerequisites: Docker with Compose v2, about 8 GB of free RAM (the stack runs nine containers).
-
Clone the repository and start the stack:
Fenêtre de terminal git clone https://forge.erythix.tech/labstraining/labs/VMLLM.gitcd VMLLMdocker compose up -d --builddocker compose ps # check that all services are running -
The simulator starts with the stack (
llm-simulatorservice, 10 requests per second, 3 models) and exposes its metrics on port 9100. Check that it produces series:Fenêtre de terminal curl -s http://localhost:9100/metrics | grep llm_requests_total | head -
Check ingestion in vmui (
http://localhost:8428/vmui) with the query:sum by (model) (rate(llm_requests_total[1m])) * 60 -
Open Grafana (
http://localhost:3000). Dashboards are provisioned at startup: open the production dashboard (uidllm-prod), defined ingrafana/dashboards/dashboard-llm-prod.json.
Lab 2: write five MetricsQL queries
Section titled “Lab 2: write five MetricsQL queries”Goal: using the simulator’s metrics, write the five core production queries, validate them in vmui, then add them to a Grafana dashboard.
Estimated time: 40 minutes. Level: intermediate. Solution: solutions/lab2-queries.md in the repository.
- Total cost by model over the last 24 hours.
- TTFT P95 by provider (
histogram_quantile). - RAG retrieval score anomaly (z-score against a 7-day baseline, see module 7; the baseline needs seven days of history, so on a fresh stack the z-score is partial).
- Top 5 tenants by cost (
topk+increase). - GPU efficiency: output tokens per joule (division between metrics).
Lab 3: configure a production alert
Section titled “Lab 3: configure a production alert”Goal: configure the RAGIndexDrift alert in vmalert, test it by injecting a score degradation, check the notification in a test Slack channel.
Estimated time: 25 minutes, plus the alert’s 15-minute for:. Level: intermediate.
- Adapt the repository’s
alerts/llm-quality.ymlandalertmanager.ymlwith your test Slack webhook URL. - Reload the rules with a GET request:
curl http://localhost:8880/-/reload(ordocker compose restart vmalert). Reload Alertmanager withdocker compose restart alertmanager. - Inject a score degradation:
make lab3-drift(orbash scripts/triggers/01_rag_index_drift.sh). The script stops the nominal simulator and starts one with--score-drift 0.5, which halves the scores. - In the vmalert UI (
http://localhost:8880), check that the alert goes to pending then fires after 15 minutes. - Confirm receipt in Slack, then restore the nominal simulator:
make lab3-restore(orbash scripts/triggers/_restore.sh).
Lab 4: recording rules and performance
Section titled “Lab 4: recording rules and performance”Goal: add the essential recording rules and compare query response times with and without them through the VictoriaMetrics API.
Estimated time: 20 minutes. Level: advanced. Solution: solutions/lab4-recording-rules.md in the repository (ignore its speedup figures, see the note on the VMLLM lab).
-
Copy the rules of module 10 into
alerts/recording-rules.yml(the repository contains a similar version), then reload vmalert:curl http://localhost:8880/-/reload. -
Wait two or three evaluation intervals, then check in vmui that the
llm:request_duration_ms:p99_5mseries exists. -
Time the full query:
Fenêtre de terminal time curl -sG 'http://localhost:8428/api/v1/query' \--data-urlencode 'query=histogram_quantile(0.99, sum by (le, model, pipeline_id) (rate(llm_request_duration_ms_bucket[5m])))' \> /dev/null -
Time the pre-computed series:
Fenêtre de terminal time curl -sG 'http://localhost:8428/api/v1/query' \--data-urlencode 'query=llm:request_duration_ms:p99_5m' \> /dev/null -
Repeat each measurement about ten times and compare the medians. Do the same with
increase(llm_cost_usd_total[24h])andllm:cost_usd:increase24h. -
Interpret: the simulator’s cardinality is low (a few hundred series), so the absolute gap will be small. It grows with the number of series the full query reads; this is the ratio to measure on your own instance before rolling the rules out widely.
Appendix A: production go-live checklist
Section titled “Appendix A: production go-live checklist”Before go-live
Section titled “Before go-live”- VictoriaMetrics is deployed with
-storageDataPathon a persistent SSD. - Retention is set:
-retentionPeriodon each instance, and-retentionFilteronly if you use the Enterprise version (otherwise, separate instances per retention). - vmbackup is configured to S3-compatible storage (restore test passed).
- vmagent or the OTel Collector ingests the four metric categories (quality, latency, cost, GPU).
- The six production vmalert rules are loaded and tested (test alert received in Slack or Teams).
- At least two recording rules are active for costs.
- The production Grafana dashboard is imported and template variables work.
- vmauth is configured with authentication (basic auth or bearer token; accepting mTLS connections is a vmauth Enterprise feature), never unauthenticated access.
- Total cardinality is measured and documented (
/api/v1/status/tsdb). - Risky labels (
trace_id,session_id) are dropped by relabelling. - A
VictoriaMetricsDownself-monitoring alert is active. - An incident runbook exists (who to call, how to restore, how to scale).
- Read-only Grafana users exist for non-technical roles (management, finance).
After go-live (D+7, D+30)
Section titled “After go-live (D+7, D+30)”- Cardinality has stabilized (no linear growth).
- Disk consumption is measured and projected to 90 days.
- At least one real alert has been received and routing confirmed.
- P99 latency of Grafana queries is measured and compared with the target you set (for example under 1 s).
- A vmbackup restore test has been run on a test instance.
- Alert thresholds have been reviewed with the teams (false positives?).
- Custom MetricsQL queries used by the teams are documented.
Appendix B: MetricsQL cheat sheet for LLMs
Section titled “Appendix B: MetricsQL cheat sheet for LLMs”Cost and tokens
Section titled “Cost and tokens”# Hourly cost per modelsum by (model) (increase(llm_cost_usd_total[1h]))
# Projected daily costsum(rate(llm_cost_usd_total[30m])) * 86400
# Top 10 tenants over 24htopk(10, sum by (tenant_id) (increase(llm_cost_usd_total[24h])))
# Average cost per request (recording rule recommended)sum(rate(llm_cost_usd_total[5m])) / sum(rate(llm_requests_total[5m]))Latency and performance
Section titled “Latency and performance”# P99 end-to-end latency per modelhistogram_quantile(0.99, sum by (le, model) (rate(llm_request_duration_ms_bucket[5m])))
# TTFT P95 (streaming)histogram_quantile(0.95, sum by (le, model) (rate(llm_ttft_ms_bucket[5m])))
# Global error ratesum(rate(llm_errors_total[5m])) / sum(rate(llm_requests_total[5m]))
# Output token throughput per model (tokens/s)sum by (model) (rate(llm_tokens_total{type='output'}[1m]))RAG quality
Section titled “RAG quality”# Retrieval score z-score against a 7-day baseline# (replaces anomaly_score(), which is not a MetricsQL function)( avg by (index_id) (rag_retrieval_score_avg) - avg_over_time(avg by (index_id) (rag_retrieval_score_avg)[7d:1h]))/ stddev_over_time(avg by (index_id) (rag_retrieval_score_avg)[7d:1h])
# If you run vmanomaly (Enterprise): query the metric it writesanomaly_score > 1
# P95 context pressure per pipelinehistogram_quantile(0.95, sum by (le, pipeline_id) (rate(rag_context_pressure_ratio_bucket[5m])))
# Share of dropped chunks (requires a rag_chunks_retrieved_total counter)sum by (pipeline_id) (rate(rag_chunks_dropped_total[5m])) / sum by (pipeline_id) (rate(rag_chunks_retrieved_total[5m]))GPU and infrastructure
Section titled “GPU and infrastructure”# Average GPU utilization per node and GPUavg by (node, gpu) (DCGM_FI_DEV_GPU_UTIL)
# VRAM saturation (alert before OOM)DCGM_FI_DEV_FB_USED / (DCGM_FI_DEV_FB_USED + DCGM_FI_DEV_FB_FREE)
# Output tokens per joule (global; DCGM series have no LLM model label)sum(rate(llm_tokens_total{type='output'}[5m])) / sum(DCGM_FI_DEV_POWER_USAGE)
# vLLM queuevllm:num_requests_waitingAdvanced patterns
Section titled “Advanced patterns”# Robust median (ignores spikes)median_over_time(rag_retrieval_score_avg[30m])
# Open, high, low, close of RAG scores (candlestick visualization)rollup_candlestick(rag_retrieval_score_avg[1h])
# Near-zero traffic: the slope of a counter is a per-second rateabs(deriv(llm_requests_total[10m])) < 0.01
# Week-over-week comparisonsum(rate(llm_cost_usd_total[1h])) / sum(rate(llm_cost_usd_total[1h] offset 7d))References
Section titled “References”Official VictoriaMetrics documentation
Section titled “Official VictoriaMetrics documentation”| Resource | URL |
|---|---|
| VictoriaMetrics documentation | docs.victoriametrics.com |
| MetricsQL reference | docs.victoriametrics.com/metricsql |
| vmagent | docs.victoriametrics.com/vmagent |
| vmalert | docs.victoriametrics.com/vmalert |
| vmauth | docs.victoriametrics.com/vmauth |
| vmanomaly | docs.victoriametrics.com/anomaly-detection |
| VictoriaMetrics source code | github.com/VictoriaMetrics/VictoriaMetrics |
Integrations and tools
Section titled “Integrations and tools”| Tool | URL |
|---|---|
| OTel Collector Contrib | github.com/open-telemetry/opentelemetry-collector-contrib |
| DCGM exporter (GPU) | github.com/NVIDIA/dcgm-exporter |
| vLLM metrics | docs.vllm.ai (section “Metrics”) |
| LiteLLM proxy metrics | docs.litellm.ai/docs/proxy/prometheus |
| Grafana VictoriaMetrics plugin | grafana.com/grafana/plugins/victoriametrics-metrics-datasource |
| prometheus_client for Python | github.com/prometheus/client_python |