People
Simulator: alert fatigue
For: engineers and SREs · team managersPrerequisites: Basic notions of on-call and alerting.
An alert that requires no action is not free: it interrupts, it wakes people up, and above all it teaches the team to stop reacting. What is called alert fatigue can be measured: how many interruptions for the on-call person, how many broken nights, how many hours lost, and what share of all that was just noise.
Parameters
Sustainability thresholds (adjustable)
Result
| Criterion | Value | Threshold | Status |
|---|
- Actionable
- Noise (non-actionable)
- Threshold
- Outside business hours (shaded)
Model based on averages: weekly rotation, one on-call person receives every alert of their week; business-hours alerts spread over the 5 weekdays, the others over the remaining 9 periods of 12 h (nights and weekend). Each alert counts as an interruption, even when several alerts belong to the same incident. The guideline of 2 incidents per 12 h shift and that of a rotation of at least 8 people are commonly cited from Google’s Site Reliability Engineering book; they are guidelines, not standards.
Things to try
Section titled “Things to try”- The starting point: 40 alerts per week of which 30% are actionable, a rotation of 6 people. All four criteria are exceeded and noise costs around €36,000 a year, not counting what it causes the team to miss.
- Remove the noise: go to 12 alerts per week, 100% actionable (keep only alerts that require action). The load per 12 h period drops below the threshold and the cost of noise falls to zero, but 2.4 wake-ups per week remain.
- Grow the rotation to 8 people: the on-call week is no lighter, but everyone goes through it less often (6.5 weeks a year instead of 8.7). It is an organizational lever, not a cure for noise.
- Defer what can wait: bring the off-hours share down to 25%, for example by turning non-urgent alerts into tickets handled the next day. Wake-ups drop below the threshold and on-call becomes sustainable.
The rule to remember
Section titled “The rule to remember”Alert volume is fixed at the source: an alert should only wake someone up if it demands immediate human action. The guidelines commonly cited from Google’s Site Reliability Engineering book (at most two incidents per 12-hour on-call shift, a rotation of at least eight people for a single-site team) are orders of magnitude to adapt, not standards; what matters is to choose thresholds explicitly, measure them and review them with the team.
Going further: the people dimension, burn rate alerts that cut needless wake-ups, organization and alert ownership and the bridges between dimensions.