Skip to content

TechnicalPractitioner

MetricsQL queries and Grafana dashboards

For: engineers and SREsPrerequisites: Have read lesson 3 of the course and know PromQL basics.

7.1 MetricsQL features useful for LLM observability

Section titled “7.1 MetricsQL features useful for LLM observability”

MetricsQL is largely compatible with PromQL: most PromQL queries run unchanged. A few extensions are particularly useful here:

FeatureUse for LLM workloads
rate() and increase() without extrapolationCost totals closer to the real values on sparse counters
Optional lookbehind window (rate(x) uses the step automatically)Simpler dashboard queries
median_over_time, quantile_over_timeRobust statistics on scores and token counts
rollup_candlestickOpen, high, low, close of a series (for example RAG scores)
outliers_madSpots series that deviate from the group (for example one model among others)
WITH templatesReuse sub-expressions in long queries

Hourly LLM cost by model (USD):

sum by (model, provider) (increase(llm_cost_usd_total[1h]))

Daily cost projection (extrapolated from the last 30 minutes):

sum by (model) (rate(llm_cost_usd_total[30m])) * 86400

Top 10 tenants by cost (last 24 hours):

topk(10, sum by (tenant_id) (increase(llm_cost_usd_total[24h])))

Output/input token ratio (detect outlier requests):

sum(rate(llm_tokens_total{type='output'}[5m]))
/
sum(rate(llm_tokens_total{type='input'}[5m]))

Retrieval score anomaly: index drift detection

Section titled “Retrieval score anomaly: index drift detection”

Z-score of the retrieval score against a 7-day baseline (hourly resolution):

(
avg by (index_id) (rag_retrieval_score_avg)
- avg_over_time(avg by (index_id) (rag_retrieval_score_avg)[7d:1h])
)
/
stddev_over_time(avg by (index_id) (rag_retrieval_score_avg)[7d:1h])

A value close to 0 means “as usual”. A strongly negative value (for example below -3, to be calibrated on your history) signals a drop in retrieval quality. Two limits: the baseline needs seven days of history, and a near-zero standard deviation (very stable score) makes the ratio unstable. MetricsQL’s outliers_mad is a complementary tool when you want to compare indexes with one another rather than with their own past.

Context pressure P95 by pipeline:

histogram_quantile(0.95,
sum by (le, pipeline_id) (
rate(rag_context_pressure_ratio_bucket[10m])
)
)

Dropped chunks per 100 requests:

sum by (pipeline_id) (rate(rag_chunks_dropped_total[5m]))
/
sum by (pipeline_id) (rate(llm_requests_total[5m]))
* 100

End-to-end P99 latency by model:

histogram_quantile(0.99,
sum by (le, model, provider) (
rate(llm_request_duration_ms_bucket[5m])
)
)

TTFT P95, time to first token (critical for streaming):

histogram_quantile(0.95,
sum by (le, model) (rate(llm_ttft_ms_bucket[5m]))
)

GPU energy efficiency, output tokens per joule:

sum(rate(llm_tokens_total{type='output'}[5m]))
/
sum(DCGM_FI_DEV_POWER_USAGE)

Module 8: Grafana dashboards, structure and panels

Section titled “Module 8: Grafana dashboards, structure and panels”

A typical LLM dashboard is organized in four functional rows matching the four dimensions of the taxonomy: service overview (traffic, latency, errors), RAG quality, cost and tokens, GPU infrastructure.

# Auto-provisioned datasource (grafana/provisioning/datasources/vm.yaml)
apiVersion: 1
datasources:
- name: VictoriaMetrics
type: victoriametrics-metrics-datasource
url: http://victoriametrics:8428
access: proxy
isDefault: true
jsonData:
timeInterval: '15s'
# Variable 'model': list of active models
label_values(llm_requests_total, model)
# Variable 'provider'
label_values(llm_requests_total{model=~'$model'}, provider)
# Variable 'pipeline_id'
label_values(rag_retrieval_score_avg, pipeline_id)
MetricPanel typeKey options
Requests/min (SLA)Stat + sparklineGreen above 0, orange if drift above 20%
P99 latencyStat + thresholdAbsolute thresholds per contractual SLO
Error rateStat + thresholdRed above 1%, orange above 0.5%
Retrieval score, 7 daysTime seriesScore line, 0.70 threshold, z-score as a second query
Context pressure P95Time series + thresholdRed zone above 0.85
Cumulative cost by modelStacked area time seriesOne color per model, USD unit
Token distributionHeatmap or histogramShows outliers (abnormally long requests)
GPU utilizationGauge + thresholdOrange at 80%, red at 95%
vLLM queue depthTime seriesAlert if more than 10 requests waiting
Retrieval score z-scoreTime seriesThreshold to calibrate (for example -3), or anomaly_score if you run vmanomaly