Skip to content

TechnicalPractitioner

Labs, go-live checklist and MetricsQL cheat sheet

For: engineers and SREsPrerequisites: Have read lessons 1 to 5 of the course and have Docker with Compose v2.

Lab 1: deploy the stack and ingest first metrics

Section titled “Lab 1: deploy the stack and ingest first metrics”

Goal: deploy the complete Docker Compose stack (VictoriaMetrics, OTel Collector, vmagent, vmalert, Alertmanager, Grafana, simulator), check the metrics in vmui and Grafana. Estimated time: 30 minutes. Level: beginner. Prerequisites: Docker with Compose v2, about 8 GB of free RAM (the stack runs nine containers).

  1. Clone the repository and start the stack:

    Fenêtre de terminal
    git clone https://forge.erythix.tech/labstraining/labs/VMLLM.git
    cd VMLLM
    docker compose up -d --build
    docker compose ps # check that all services are running
  2. The simulator starts with the stack (llm-simulator service, 10 requests per second, 3 models) and exposes its metrics on port 9100. Check that it produces series:

    Fenêtre de terminal
    curl -s http://localhost:9100/metrics | grep llm_requests_total | head
  3. Check ingestion in vmui (http://localhost:8428/vmui) with the query:

    sum by (model) (rate(llm_requests_total[1m])) * 60
  4. Open Grafana (http://localhost:3000). Dashboards are provisioned at startup: open the production dashboard (uid llm-prod), defined in grafana/dashboards/dashboard-llm-prod.json.

Goal: using the simulator’s metrics, write the five core production queries, validate them in vmui, then add them to a Grafana dashboard. Estimated time: 40 minutes. Level: intermediate. Solution: solutions/lab2-queries.md in the repository.

  1. Total cost by model over the last 24 hours.
  2. TTFT P95 by provider (histogram_quantile).
  3. RAG retrieval score anomaly (z-score against a 7-day baseline, see module 7; the baseline needs seven days of history, so on a fresh stack the z-score is partial).
  4. Top 5 tenants by cost (topk + increase).
  5. GPU efficiency: output tokens per joule (division between metrics).

Goal: configure the RAGIndexDrift alert in vmalert, test it by injecting a score degradation, check the notification in a test Slack channel. Estimated time: 25 minutes, plus the alert’s 15-minute for:. Level: intermediate.

  1. Adapt the repository’s alerts/llm-quality.yml and alertmanager.yml with your test Slack webhook URL.
  2. Reload the rules with a GET request: curl http://localhost:8880/-/reload (or docker compose restart vmalert). Reload Alertmanager with docker compose restart alertmanager.
  3. Inject a score degradation: make lab3-drift (or bash scripts/triggers/01_rag_index_drift.sh). The script stops the nominal simulator and starts one with --score-drift 0.5, which halves the scores.
  4. In the vmalert UI (http://localhost:8880), check that the alert goes to pending then fires after 15 minutes.
  5. Confirm receipt in Slack, then restore the nominal simulator: make lab3-restore (or bash scripts/triggers/_restore.sh).

Goal: add the essential recording rules and compare query response times with and without them through the VictoriaMetrics API. Estimated time: 20 minutes. Level: advanced. Solution: solutions/lab4-recording-rules.md in the repository (ignore its speedup figures, see the note on the VMLLM lab).

  1. Copy the rules of module 10 into alerts/recording-rules.yml (the repository contains a similar version), then reload vmalert: curl http://localhost:8880/-/reload.

  2. Wait two or three evaluation intervals, then check in vmui that the llm:request_duration_ms:p99_5m series exists.

  3. Time the full query:

    Fenêtre de terminal
    time curl -sG 'http://localhost:8428/api/v1/query' \
    --data-urlencode 'query=histogram_quantile(0.99, sum by (le, model, pipeline_id) (rate(llm_request_duration_ms_bucket[5m])))' \
    > /dev/null
  4. Time the pre-computed series:

    Fenêtre de terminal
    time curl -sG 'http://localhost:8428/api/v1/query' \
    --data-urlencode 'query=llm:request_duration_ms:p99_5m' \
    > /dev/null
  5. Repeat each measurement about ten times and compare the medians. Do the same with increase(llm_cost_usd_total[24h]) and llm:cost_usd:increase24h.

  6. Interpret: the simulator’s cardinality is low (a few hundred series), so the absolute gap will be small. It grows with the number of series the full query reads; this is the ratio to measure on your own instance before rolling the rules out widely.

  • VictoriaMetrics is deployed with -storageDataPath on a persistent SSD.
  • Retention is set: -retentionPeriod on each instance, and -retentionFilter only if you use the Enterprise version (otherwise, separate instances per retention).
  • vmbackup is configured to S3-compatible storage (restore test passed).
  • vmagent or the OTel Collector ingests the four metric categories (quality, latency, cost, GPU).
  • The six production vmalert rules are loaded and tested (test alert received in Slack or Teams).
  • At least two recording rules are active for costs.
  • The production Grafana dashboard is imported and template variables work.
  • vmauth is configured with authentication (basic auth or bearer token; accepting mTLS connections is a vmauth Enterprise feature), never unauthenticated access.
  • Total cardinality is measured and documented (/api/v1/status/tsdb).
  • Risky labels (trace_id, session_id) are dropped by relabelling.
  • A VictoriaMetricsDown self-monitoring alert is active.
  • An incident runbook exists (who to call, how to restore, how to scale).
  • Read-only Grafana users exist for non-technical roles (management, finance).
  • Cardinality has stabilized (no linear growth).
  • Disk consumption is measured and projected to 90 days.
  • At least one real alert has been received and routing confirmed.
  • P99 latency of Grafana queries is measured and compared with the target you set (for example under 1 s).
  • A vmbackup restore test has been run on a test instance.
  • Alert thresholds have been reviewed with the teams (false positives?).
  • Custom MetricsQL queries used by the teams are documented.

Appendix B: MetricsQL cheat sheet for LLMs

Section titled “Appendix B: MetricsQL cheat sheet for LLMs”
# Hourly cost per model
sum by (model) (increase(llm_cost_usd_total[1h]))
# Projected daily cost
sum(rate(llm_cost_usd_total[30m])) * 86400
# Top 10 tenants over 24h
topk(10, sum by (tenant_id) (increase(llm_cost_usd_total[24h])))
# Average cost per request (recording rule recommended)
sum(rate(llm_cost_usd_total[5m])) / sum(rate(llm_requests_total[5m]))
# P99 end-to-end latency per model
histogram_quantile(0.99, sum by (le, model) (rate(llm_request_duration_ms_bucket[5m])))
# TTFT P95 (streaming)
histogram_quantile(0.95, sum by (le, model) (rate(llm_ttft_ms_bucket[5m])))
# Global error rate
sum(rate(llm_errors_total[5m])) / sum(rate(llm_requests_total[5m]))
# Output token throughput per model (tokens/s)
sum by (model) (rate(llm_tokens_total{type='output'}[1m]))
# Retrieval score z-score against a 7-day baseline
# (replaces anomaly_score(), which is not a MetricsQL function)
(
avg by (index_id) (rag_retrieval_score_avg)
- avg_over_time(avg by (index_id) (rag_retrieval_score_avg)[7d:1h])
)
/ stddev_over_time(avg by (index_id) (rag_retrieval_score_avg)[7d:1h])
# If you run vmanomaly (Enterprise): query the metric it writes
anomaly_score > 1
# P95 context pressure per pipeline
histogram_quantile(0.95, sum by (le, pipeline_id) (rate(rag_context_pressure_ratio_bucket[5m])))
# Share of dropped chunks (requires a rag_chunks_retrieved_total counter)
sum by (pipeline_id) (rate(rag_chunks_dropped_total[5m]))
/ sum by (pipeline_id) (rate(rag_chunks_retrieved_total[5m]))
# Average GPU utilization per node and GPU
avg by (node, gpu) (DCGM_FI_DEV_GPU_UTIL)
# VRAM saturation (alert before OOM)
DCGM_FI_DEV_FB_USED / (DCGM_FI_DEV_FB_USED + DCGM_FI_DEV_FB_FREE)
# Output tokens per joule (global; DCGM series have no LLM model label)
sum(rate(llm_tokens_total{type='output'}[5m])) / sum(DCGM_FI_DEV_POWER_USAGE)
# vLLM queue
vllm:num_requests_waiting
# Robust median (ignores spikes)
median_over_time(rag_retrieval_score_avg[30m])
# Open, high, low, close of RAG scores (candlestick visualization)
rollup_candlestick(rag_retrieval_score_avg[1h])
# Near-zero traffic: the slope of a counter is a per-second rate
abs(deriv(llm_requests_total[10m])) < 0.01
# Week-over-week comparison
sum(rate(llm_cost_usd_total[1h]))
/ sum(rate(llm_cost_usd_total[1h] offset 7d))
ResourceURL
VictoriaMetrics documentationdocs.victoriametrics.com
MetricsQL referencedocs.victoriametrics.com/metricsql
vmagentdocs.victoriametrics.com/vmagent
vmalertdocs.victoriametrics.com/vmalert
vmauthdocs.victoriametrics.com/vmauth
vmanomalydocs.victoriametrics.com/anomaly-detection
VictoriaMetrics source codegithub.com/VictoriaMetrics/VictoriaMetrics
ToolURL
OTel Collector Contribgithub.com/open-telemetry/opentelemetry-collector-contrib
DCGM exporter (GPU)github.com/NVIDIA/dcgm-exporter
vLLM metricsdocs.vllm.ai (section “Metrics”)
LiteLLM proxy metricsdocs.litellm.ai/docs/proxy/prometheus
Grafana VictoriaMetrics plugingrafana.com/grafana/plugins/victoriametrics-metrics-datasource
prometheus_client for Pythongithub.com/prometheus/client_python