Skip to content

TechnicalPractitioner

Alerts, recording rules and SLOs

For: engineers and SREs · architectsPrerequisites: Have read lesson 4 of the course on MetricsQL.

alerts/llm-quality.yml
groups:
- name: llm_quality
interval: 1m
rules:
# RAG index drift
- alert: RAGIndexDrift
expr: |
avg by (index_id, pipeline_id) (
avg_over_time(rag_retrieval_score_avg[1h])
) < 0.70
for: 15m
labels: { severity: warning, team: ai-platform }
annotations:
summary: 'RAG index drift on {{ $labels.index_id }}'
description: |
Mean retrieval score = {{ $value | printf "%.2f" }} < threshold 0.70
for 15m on pipeline {{ $labels.pipeline_id }}.
Action: check for document updates without re-indexing.
runbook_url: 'https://wiki.internal/runbooks/rag-index-drift'
# Critical context pressure
- alert: RAGContextPressureHigh
expr: |
histogram_quantile(0.95,
sum by (le, pipeline_id) (
rate(rag_context_pressure_ratio_bucket[10m])
)
) > 0.85
for: 5m
labels: { severity: warning, team: ai-platform }
annotations:
summary: 'High context pressure P95: {{ $labels.pipeline_id }}'
description: 'P95 = {{ $value | printf "%.2f" }}. Reduce the system prompt or increase the context window.'
# High LLM error rate
- alert: LLMHighErrorRate
expr: |
sum by (model, provider) (rate(llm_errors_total[5m]))
/
sum by (model, provider) (rate(llm_requests_total[5m]))
> 0.01
for: 3m
labels: { severity: critical, team: ai-platform }
annotations:
summary: 'Error rate > 1% on {{ $labels.model }} / {{ $labels.provider }}'
description: 'Current rate: {{ $value | humanizePercentage }}. Check provider availability.'
# Latency SLO breach
- alert: LLMLatencySLOBreach
expr: |
histogram_quantile(0.99,
sum by (le, model, pipeline_id) (
rate(llm_request_duration_ms_bucket[5m])
)
) > 5000
for: 5m
labels: { severity: critical, team: ai-platform }
annotations:
summary: 'P99 latency SLO breach on {{ $labels.model }}'
description: 'P99 = {{ $value | printf "%.0f" }}ms > SLO 5000ms. Pipeline {{ $labels.pipeline_id }}'
# GPU memory saturation
- alert: GPUMemorySaturation
expr: DCGM_FI_DEV_FB_USED / (DCGM_FI_DEV_FB_USED + DCGM_FI_DEV_FB_FREE) * 100 > 90
for: 2m
labels: { severity: warning, team: mlops }
annotations:
summary: 'GPU {{ $labels.gpu }} memory saturated ({{ $value | printf "%.0f" }}%)'
description: 'OOM risk. Reduce batch size or model context window.'
# Daily LLM budget (example budget: 100 USD per day, replace with yours).
# llm_cost_usd_total exists only if a price grid is supplied (module 6.3).
- alert: LLMDailyBudgetWarning
expr: |
sum by (tenant_id) (
increase(llm_cost_usd_total[24h])
) > 80
labels: { severity: warning, team: billing }
annotations:
summary: 'LLM budget at 80% for tenant {{ $labels.tenant_id }}'
description: '24h cost = ${{ $value | printf "%.2f" }}. Threshold: $100/day.'
alertmanager.yml
global:
slack_api_url: 'https://hooks.slack.com/services/REPLACE_ME'
route:
group_by: ['alertname', 'model', 'pipeline_id']
group_wait: 30s
repeat_interval: 12h
receiver: 'default'
routes:
- matchers: ['severity="critical"']
receiver: pagerduty
- matchers: ['team="ai-platform"']
receiver: slack-ai-platform
- matchers: ['team="billing"']
receiver: slack-billing
receivers:
- name: 'default'
slack_configs:
- channel: '#llm-ops-alerts'
title: '{{ .CommonAnnotations.summary }}'
text: '{{ .CommonAnnotations.description }}'
- name: 'slack-ai-platform'
slack_configs:
- channel: '#ai-platform-alerts'
- name: 'slack-billing'
slack_configs:
- channel: '#billing-alerts'
- name: 'pagerduty'
pagerduty_configs:
- routing_key: 'YOUR_PAGERDUTY_KEY'

Module 10: recording rules, pre-aggregating cost metrics

Section titled “Module 10: recording rules, pre-aggregating cost metrics”

Cost metrics are queried very frequently in dashboards. Recalculating increase(llm_cost_usd_total[24h]) on every refresh is CPU-intensive. Recording rules pre-compute these aggregations in the background: the dashboard then reads one series per label combination instead of recomputing the aggregation over all raw series. The gain depends on the number of series and on the queried window; measure it on your instance (this is what Lab 4 is for).

alerts/recording-rules.yml
groups:
- name: llm_recording_rules
interval: 1m
rules:
- record: llm:cost_usd:increase1h
expr: |
sum by (model, provider, tenant_id) (
increase(llm_cost_usd_total[1h])
)
- record: llm:cost_usd:increase24h
expr: |
sum by (model, provider, tenant_id) (
increase(llm_cost_usd_total[24h])
)
- record: llm:error_rate:ratio5m
expr: |
sum by (model, provider) (rate(llm_errors_total[5m]))
/
sum by (model, provider) (rate(llm_requests_total[5m]))
- record: llm:request_duration_ms:p99_5m
expr: |
histogram_quantile(0.99,
sum by (le, model, pipeline_id) (
rate(llm_request_duration_ms_bucket[5m])
)
)
- record: llm:ttft_ms:p95_5m
expr: |
histogram_quantile(0.95,
sum by (le, model) (rate(llm_ttft_ms_bucket[5m]))
)
- record: rag:retrieval_score:avg30m
expr: |
avg by (index_id, pipeline_id) (
avg_over_time(rag_retrieval_score_avg[30m])
)
- record: rag:context_pressure:p95_10m
expr: |
histogram_quantile(0.95,
sum by (le, pipeline_id) (
rate(rag_context_pressure_ratio_bucket[10m])
)
)
- record: gpu:memory_utilization:ratio
expr: DCGM_FI_DEV_FB_USED / (DCGM_FI_DEV_FB_USED + DCGM_FI_DEV_FB_FREE)

Using recording rules in alerts and dashboards

Section titled “Using recording rules in alerts and dashboards”

Once defined, use the recorded series instead of the full query. Instead of:

histogram_quantile(0.99, sum by (le, model, pipeline_id) (rate(llm_request_duration_ms_bucket[5m])))

use:

llm:request_duration_ms:p99_5m

The query is much faster. The result is equivalent, at the resolution of the rule’s evaluation interval (here one minute), and with the aggregation labels chosen in the rule.

A Service Level Objective (SLO) translates a business requirement into a measurable target. For LLMs, classic SLOs (availability, latency) still apply but are not enough: add quality and cost.

CategoryIndicator (SLI)SLO (example to adapt)Window
Availability1 - rate(llm_errors_total[5m]) / rate(llm_requests_total[5m])99.5%30 days
Interactive latencyhistogram_quantile(0.95, ...llm_ttft_ms...)< 1,200 ms7 days
Batch latencyhistogram_quantile(0.99, ...llm_request_duration_ms...)< 8,000 ms30 days
Retrieval qualityavg(rag_retrieval_score_avg)> 0.7824 h
Consumption per requestsum(increase(llm_tokens_total[1h])) / sum(increase(llm_requests_total[1h]))< 2,000 tokens (example to adapt)Calendar month
Cost per request (with a price grid)sum(increase(llm_cost_usd_total[1h])) / sum(increase(llm_requests_total[1h]))your target, in the metric’s currencyCalendar month

The error budget is the complement of the SLO to 100%. For 99.5% availability over 30 days, the error budget is 0.5%, that is 3 h 36 min of tolerated unavailability per month.

# Error budget consumed over the SLO window (in %)
(
1
-
(
sum(increase(llm_requests_total[30d]))
-
sum(increase(llm_errors_total[30d]))
)
/
sum(increase(llm_requests_total[30d]))
) / 0.005 * 100

11.3 Burn rate alerts (Google SRE pattern)

Section titled “11.3 Burn rate alerts (Google SRE pattern)”

Alert when the error budget is consumed too fast. Each alert combines a long window and a short window: the long window confirms the trend, the short one makes the alert stop quickly once the problem is fixed.

- alert: LLMFastBurn
expr: |
(
sum(rate(llm_errors_total[1h])) / sum(rate(llm_requests_total[1h]))
) > (14.4 * 0.005)
and
(
sum(rate(llm_errors_total[5m])) / sum(rate(llm_requests_total[5m]))
) > (14.4 * 0.005)
for: 2m
labels:
severity: critical
annotations:
summary: "Error budget burn rate 14.4x normal (budget exhausted in about 2 days)"
- alert: LLMSlowBurn
expr: |
(
sum(rate(llm_errors_total[6h])) / sum(rate(llm_requests_total[6h]))
) > (6 * 0.005)
and
(
sum(rate(llm_errors_total[30m])) / sum(rate(llm_requests_total[30m]))
) > (6 * 0.005)
for: 15m
labels:
severity: warning
annotations:
summary: "Error budget burn rate 6x normal (budget exhausted in 5 days)"

Revised on 2 October 2026: prices removed; the numeric cost per request target is replaced by a tokens per request target, the cost target being left to your price grid.