Technical
Alerts, recording rules and SLOs
For: engineers and SREs · architectsPrerequisites: Have read lesson 4 of the course on MetricsQL.
Module 9: vmalert, production alert rules
Section titled “Module 9: vmalert, production alert rules”9.1 Complete production ruleset
Section titled “9.1 Complete production ruleset”groups: - name: llm_quality interval: 1m rules:
# RAG index drift - alert: RAGIndexDrift expr: | avg by (index_id, pipeline_id) ( avg_over_time(rag_retrieval_score_avg[1h]) ) < 0.70 for: 15m labels: { severity: warning, team: ai-platform } annotations: summary: 'RAG index drift on {{ $labels.index_id }}' description: | Mean retrieval score = {{ $value | printf "%.2f" }} < threshold 0.70 for 15m on pipeline {{ $labels.pipeline_id }}. Action: check for document updates without re-indexing. runbook_url: 'https://wiki.internal/runbooks/rag-index-drift'
# Critical context pressure - alert: RAGContextPressureHigh expr: | histogram_quantile(0.95, sum by (le, pipeline_id) ( rate(rag_context_pressure_ratio_bucket[10m]) ) ) > 0.85 for: 5m labels: { severity: warning, team: ai-platform } annotations: summary: 'High context pressure P95: {{ $labels.pipeline_id }}' description: 'P95 = {{ $value | printf "%.2f" }}. Reduce the system prompt or increase the context window.'
# High LLM error rate - alert: LLMHighErrorRate expr: | sum by (model, provider) (rate(llm_errors_total[5m])) / sum by (model, provider) (rate(llm_requests_total[5m])) > 0.01 for: 3m labels: { severity: critical, team: ai-platform } annotations: summary: 'Error rate > 1% on {{ $labels.model }} / {{ $labels.provider }}' description: 'Current rate: {{ $value | humanizePercentage }}. Check provider availability.'
# Latency SLO breach - alert: LLMLatencySLOBreach expr: | histogram_quantile(0.99, sum by (le, model, pipeline_id) ( rate(llm_request_duration_ms_bucket[5m]) ) ) > 5000 for: 5m labels: { severity: critical, team: ai-platform } annotations: summary: 'P99 latency SLO breach on {{ $labels.model }}' description: 'P99 = {{ $value | printf "%.0f" }}ms > SLO 5000ms. Pipeline {{ $labels.pipeline_id }}'
# GPU memory saturation - alert: GPUMemorySaturation expr: DCGM_FI_DEV_FB_USED / (DCGM_FI_DEV_FB_USED + DCGM_FI_DEV_FB_FREE) * 100 > 90 for: 2m labels: { severity: warning, team: mlops } annotations: summary: 'GPU {{ $labels.gpu }} memory saturated ({{ $value | printf "%.0f" }}%)' description: 'OOM risk. Reduce batch size or model context window.'
# Daily LLM budget (example budget: 100 USD per day, replace with yours). # llm_cost_usd_total exists only if a price grid is supplied (module 6.3). - alert: LLMDailyBudgetWarning expr: | sum by (tenant_id) ( increase(llm_cost_usd_total[24h]) ) > 80 labels: { severity: warning, team: billing } annotations: summary: 'LLM budget at 80% for tenant {{ $labels.tenant_id }}' description: '24h cost = ${{ $value | printf "%.2f" }}. Threshold: $100/day.'9.2 Alertmanager routing by team
Section titled “9.2 Alertmanager routing by team”global: slack_api_url: 'https://hooks.slack.com/services/REPLACE_ME'
route: group_by: ['alertname', 'model', 'pipeline_id'] group_wait: 30s repeat_interval: 12h receiver: 'default' routes: - matchers: ['severity="critical"'] receiver: pagerduty - matchers: ['team="ai-platform"'] receiver: slack-ai-platform - matchers: ['team="billing"'] receiver: slack-billing
receivers: - name: 'default' slack_configs: - channel: '#llm-ops-alerts' title: '{{ .CommonAnnotations.summary }}' text: '{{ .CommonAnnotations.description }}' - name: 'slack-ai-platform' slack_configs: - channel: '#ai-platform-alerts' - name: 'slack-billing' slack_configs: - channel: '#billing-alerts' - name: 'pagerduty' pagerduty_configs: - routing_key: 'YOUR_PAGERDUTY_KEY'Module 10: recording rules, pre-aggregating cost metrics
Section titled “Module 10: recording rules, pre-aggregating cost metrics”Cost metrics are queried very frequently in dashboards. Recalculating increase(llm_cost_usd_total[24h]) on every refresh is CPU-intensive. Recording rules pre-compute these aggregations in the background: the dashboard then reads one series per label combination instead of recomputing the aggregation over all raw series. The gain depends on the number of series and on the queried window; measure it on your instance (this is what Lab 4 is for).
groups: - name: llm_recording_rules interval: 1m rules:
- record: llm:cost_usd:increase1h expr: | sum by (model, provider, tenant_id) ( increase(llm_cost_usd_total[1h]) )
- record: llm:cost_usd:increase24h expr: | sum by (model, provider, tenant_id) ( increase(llm_cost_usd_total[24h]) )
- record: llm:error_rate:ratio5m expr: | sum by (model, provider) (rate(llm_errors_total[5m])) / sum by (model, provider) (rate(llm_requests_total[5m]))
- record: llm:request_duration_ms:p99_5m expr: | histogram_quantile(0.99, sum by (le, model, pipeline_id) ( rate(llm_request_duration_ms_bucket[5m]) ) )
- record: llm:ttft_ms:p95_5m expr: | histogram_quantile(0.95, sum by (le, model) (rate(llm_ttft_ms_bucket[5m])) )
- record: rag:retrieval_score:avg30m expr: | avg by (index_id, pipeline_id) ( avg_over_time(rag_retrieval_score_avg[30m]) )
- record: rag:context_pressure:p95_10m expr: | histogram_quantile(0.95, sum by (le, pipeline_id) ( rate(rag_context_pressure_ratio_bucket[10m]) ) )
- record: gpu:memory_utilization:ratio expr: DCGM_FI_DEV_FB_USED / (DCGM_FI_DEV_FB_USED + DCGM_FI_DEV_FB_FREE)Using recording rules in alerts and dashboards
Section titled “Using recording rules in alerts and dashboards”Once defined, use the recorded series instead of the full query. Instead of:
histogram_quantile(0.99, sum by (le, model, pipeline_id) (rate(llm_request_duration_ms_bucket[5m])))use:
llm:request_duration_ms:p99_5mThe query is much faster. The result is equivalent, at the resolution of the rule’s evaluation interval (here one minute), and with the aggregation labels chosen in the rule.
Module 11: SLO and SLA for LLM workloads
Section titled “Module 11: SLO and SLA for LLM workloads”A Service Level Objective (SLO) translates a business requirement into a measurable target. For LLMs, classic SLOs (availability, latency) still apply but are not enough: add quality and cost.
11.1 Five-category LLM SLO model
Section titled “11.1 Five-category LLM SLO model”| Category | Indicator (SLI) | SLO (example to adapt) | Window |
|---|---|---|---|
| Availability | 1 - rate(llm_errors_total[5m]) / rate(llm_requests_total[5m]) | 99.5% | 30 days |
| Interactive latency | histogram_quantile(0.95, ...llm_ttft_ms...) | < 1,200 ms | 7 days |
| Batch latency | histogram_quantile(0.99, ...llm_request_duration_ms...) | < 8,000 ms | 30 days |
| Retrieval quality | avg(rag_retrieval_score_avg) | > 0.78 | 24 h |
| Consumption per request | sum(increase(llm_tokens_total[1h])) / sum(increase(llm_requests_total[1h])) | < 2,000 tokens (example to adapt) | Calendar month |
| Cost per request (with a price grid) | sum(increase(llm_cost_usd_total[1h])) / sum(increase(llm_requests_total[1h])) | your target, in the metric’s currency | Calendar month |
11.2 Error budget calculation
Section titled “11.2 Error budget calculation”The error budget is the complement of the SLO to 100%. For 99.5% availability over 30 days, the error budget is 0.5%, that is 3 h 36 min of tolerated unavailability per month.
# Error budget consumed over the SLO window (in %)( 1 - ( sum(increase(llm_requests_total[30d])) - sum(increase(llm_errors_total[30d])) ) / sum(increase(llm_requests_total[30d]))) / 0.005 * 10011.3 Burn rate alerts (Google SRE pattern)
Section titled “11.3 Burn rate alerts (Google SRE pattern)”Alert when the error budget is consumed too fast. Each alert combines a long window and a short window: the long window confirms the trend, the short one makes the alert stop quickly once the problem is fixed.
- alert: LLMFastBurn expr: | ( sum(rate(llm_errors_total[1h])) / sum(rate(llm_requests_total[1h])) ) > (14.4 * 0.005) and ( sum(rate(llm_errors_total[5m])) / sum(rate(llm_requests_total[5m])) ) > (14.4 * 0.005) for: 2m labels: severity: critical annotations: summary: "Error budget burn rate 14.4x normal (budget exhausted in about 2 days)"
- alert: LLMSlowBurn expr: | ( sum(rate(llm_errors_total[6h])) / sum(rate(llm_requests_total[6h])) ) > (6 * 0.005) and ( sum(rate(llm_errors_total[30m])) / sum(rate(llm_requests_total[30m])) ) > (6 * 0.005) for: 15m labels: severity: warning annotations: summary: "Error budget burn rate 6x normal (budget exhausted in 5 days)"Revised on 2 October 2026: prices removed; the numeric cost per request target is replaced by a tokens per request target, the cost target being left to your price grid.