Technical
MetricsQL queries and Grafana dashboards
For: engineers and SREsPrerequisites: Have read lesson 3 of the course and know PromQL basics.
Module 7: MetricsQL, essential queries
Section titled “Module 7: MetricsQL, essential queries”7.1 MetricsQL features useful for LLM observability
Section titled “7.1 MetricsQL features useful for LLM observability”MetricsQL is largely compatible with PromQL: most PromQL queries run unchanged. A few extensions are particularly useful here:
| Feature | Use for LLM workloads |
|---|---|
rate() and increase() without extrapolation | Cost totals closer to the real values on sparse counters |
Optional lookbehind window (rate(x) uses the step automatically) | Simpler dashboard queries |
median_over_time, quantile_over_time | Robust statistics on scores and token counts |
rollup_candlestick | Open, high, low, close of a series (for example RAG scores) |
outliers_mad | Spots series that deviate from the group (for example one model among others) |
WITH templates | Reuse sub-expressions in long queries |
7.2 Cost and token queries
Section titled “7.2 Cost and token queries”Hourly LLM cost by model (USD):
sum by (model, provider) (increase(llm_cost_usd_total[1h]))Daily cost projection (extrapolated from the last 30 minutes):
sum by (model) (rate(llm_cost_usd_total[30m])) * 86400Top 10 tenants by cost (last 24 hours):
topk(10, sum by (tenant_id) (increase(llm_cost_usd_total[24h])))Output/input token ratio (detect outlier requests):
sum(rate(llm_tokens_total{type='output'}[5m])) /sum(rate(llm_tokens_total{type='input'}[5m]))7.3 RAG quality queries
Section titled “7.3 RAG quality queries”Retrieval score anomaly: index drift detection
Section titled “Retrieval score anomaly: index drift detection”Z-score of the retrieval score against a 7-day baseline (hourly resolution):
( avg by (index_id) (rag_retrieval_score_avg) - avg_over_time(avg by (index_id) (rag_retrieval_score_avg)[7d:1h]))/stddev_over_time(avg by (index_id) (rag_retrieval_score_avg)[7d:1h])A value close to 0 means “as usual”. A strongly negative value (for example below -3, to be calibrated on your history) signals a drop in retrieval quality. Two limits: the baseline needs seven days of history, and a near-zero standard deviation (very stable score) makes the ratio unstable. MetricsQL’s outliers_mad is a complementary tool when you want to compare indexes with one another rather than with their own past.
Context pressure P95 by pipeline:
histogram_quantile(0.95, sum by (le, pipeline_id) ( rate(rag_context_pressure_ratio_bucket[10m]) ))Dropped chunks per 100 requests:
sum by (pipeline_id) (rate(rag_chunks_dropped_total[5m])) /sum by (pipeline_id) (rate(llm_requests_total[5m])) * 1007.4 Performance and latency queries
Section titled “7.4 Performance and latency queries”End-to-end P99 latency by model:
histogram_quantile(0.99, sum by (le, model, provider) ( rate(llm_request_duration_ms_bucket[5m]) ))TTFT P95, time to first token (critical for streaming):
histogram_quantile(0.95, sum by (le, model) (rate(llm_ttft_ms_bucket[5m])))GPU energy efficiency, output tokens per joule:
sum(rate(llm_tokens_total{type='output'}[5m])) /sum(DCGM_FI_DEV_POWER_USAGE)Module 8: Grafana dashboards, structure and panels
Section titled “Module 8: Grafana dashboards, structure and panels”A typical LLM dashboard is organized in four functional rows matching the four dimensions of the taxonomy: service overview (traffic, latency, errors), RAG quality, cost and tokens, GPU infrastructure.
8.1 Datasource configuration
Section titled “8.1 Datasource configuration”# Auto-provisioned datasource (grafana/provisioning/datasources/vm.yaml)apiVersion: 1datasources: - name: VictoriaMetrics type: victoriametrics-metrics-datasource url: http://victoriametrics:8428 access: proxy isDefault: true jsonData: timeInterval: '15s'8.2 Template variables
Section titled “8.2 Template variables”# Variable 'model': list of active modelslabel_values(llm_requests_total, model)
# Variable 'provider'label_values(llm_requests_total{model=~'$model'}, provider)
# Variable 'pipeline_id'label_values(rag_retrieval_score_avg, pipeline_id)8.3 Recommended panel types per metric
Section titled “8.3 Recommended panel types per metric”| Metric | Panel type | Key options |
|---|---|---|
| Requests/min (SLA) | Stat + sparkline | Green above 0, orange if drift above 20% |
| P99 latency | Stat + threshold | Absolute thresholds per contractual SLO |
| Error rate | Stat + threshold | Red above 1%, orange above 0.5% |
| Retrieval score, 7 days | Time series | Score line, 0.70 threshold, z-score as a second query |
| Context pressure P95 | Time series + threshold | Red zone above 0.85 |
| Cumulative cost by model | Stacked area time series | One color per model, USD unit |
| Token distribution | Heatmap or histogram | Shows outliers (abnormally long requests) |
| GPU utilization | Gauge + threshold | Orange at 80%, red at 95% |
| vLLM queue depth | Time series | Alert if more than 10 requests waiting |
| Retrieval score z-score | Time series | Threshold to calibrate (for example -3), or anomaly_score if you run vmanomaly |