Skip to content

PeoplePractitioner

Simulator: alert fatigue

For: engineers and SREs · team managersPrerequisites: Basic notions of on-call and alerting.

Reading mode

An alert that requires no action is not free: it interrupts, it wakes people up, and above all it teaches the team to stop reacting. What is called alert fatigue can be measured: how many interruptions for the on-call person, how many broken nights, how many hours lost, and what share of all that was just noise.

Parameters

Sustainability thresholds (adjustable)

Result

Sustainability
Alerts per on-call week
Night wake-ups per on-call week
Hours lost per on-call week
Noise share
Annual cost of noise
CriterionValueThresholdStatus
Expected interruptions per 12 h period over one on-call week
  • Actionable
  • Noise (non-actionable)
  • Threshold
  • Outside business hours (shaded)

Model based on averages: weekly rotation, one on-call person receives every alert of their week; business-hours alerts spread over the 5 weekdays, the others over the remaining 9 periods of 12 h (nights and weekend). Each alert counts as an interruption, even when several alerts belong to the same incident. The guideline of 2 incidents per 12 h shift and that of a rotation of at least 8 people are commonly cited from Google’s Site Reliability Engineering book; they are guidelines, not standards.

  1. The starting point: 40 alerts per week of which 30% are actionable, a rotation of 6 people. All four criteria are exceeded and noise costs around €36,000 a year, not counting what it causes the team to miss.
  2. Remove the noise: go to 12 alerts per week, 100% actionable (keep only alerts that require action). The load per 12 h period drops below the threshold and the cost of noise falls to zero, but 2.4 wake-ups per week remain.
  3. Grow the rotation to 8 people: the on-call week is no lighter, but everyone goes through it less often (6.5 weeks a year instead of 8.7). It is an organizational lever, not a cure for noise.
  4. Defer what can wait: bring the off-hours share down to 25%, for example by turning non-urgent alerts into tickets handled the next day. Wake-ups drop below the threshold and on-call becomes sustainable.

Alert volume is fixed at the source: an alert should only wake someone up if it demands immediate human action. The guidelines commonly cited from Google’s Site Reliability Engineering book (at most two incidents per 12-hour on-call shift, a rotation of at least eight people for a single-site team) are orders of magnitude to adapt, not standards; what matters is to choose thresholds explicitly, measure them and review them with the team.

Going further: the people dimension, burn rate alerts that cut needless wake-ups, organization and alert ownership and the bridges between dimensions.