Skip to content

TechnicalPractitioner

Simulator: error budget and burn rate

For: engineers and SREs · team managersPrerequisites: Basic notions of SLOs.

Reading mode

A 99.9% SLO allows 0.1% failures over the window: that is the error budget. The burn rate measures how fast you consume it: at 1 it runs out exactly at the end of the window; at 14.4 it runs out in about two days on a 30-day window.

Parameters

Result

Error budget
failed requests allowed, equivalent to
Budget left at end of window
Error budget remaining across the window

Burn rate alerts

AlertRuleFired

Rules for a 30-day SLO. Three come from the recommended table in Google's SRE Workbook, chapter "Alerting on SLOs" (14.4 over 1 h, 6 over 6 h, 1 over 3 days); the 3 over 24 h rule comes from the default windows of tools such as Sloth (10% of the budget in one day), and appears in the Workbook only in a configuration example. Simplification: only each rule's long window is evaluated, hourly; in production a short window confirms the alert and clears it once the incident is over.

  1. A hard outage: 20% errors for 1 hour. The fast alert fires within the hour: someone must be paged.
  2. A slow degradation: 0.3% errors for 5 days. No fast alert, yet the budget melts: only slow alerts see it, by design.
  3. Background noise too high: set the incident duration to 0 and the background rate to 0.08%. With no incident at all, 80% of the budget is already gone: the SLO is miscalibrated or the service needs hardening.
  4. A stricter SLO: switch to 99.99%. The budget is divided by ten: background noise alone is enough to exhaust it.

A single “error rate above X” rule is either too sensitive (it pages for a five-minute spike) or too slow (it misses a long degradation). Multi-window rules pair a burn rate factor with a duration: the faster the burn, the shorter the window and the more urgent the alert.