People
Sustainable on-call
For: engineers and SREs · team managersPrerequisites: Basic notions of alerting and SLOs.
On-call is not sustainable because the team is brave. It is sustainable because four things hold at the same time: a large enough rotation, few alerts that are all actionable, recognized compensation and regular measurement of the real load. When one is missing, the others eventually give way.
This article gathers the published benchmarks, a measurement method and the rules I recommend. It complements lesson 5 of the CIO path, which looks at the topic from the leadership side.
Published benchmarks
Section titled “Published benchmarks”Google’s Site Reliability Engineering book has a chapter on on-call (Being On-Call). It gives orders of magnitude that are widely quoted:
| Benchmark | Value proposed by the book |
|---|---|
| Share of an engineer’s time spent on call | at most 25% |
| Time spent on engineering (not operations) | at least 50% |
| Incidents per 12-hour on-call shift | at most two |
| Rotation size, single site | at least eight people |
| Rotation size, two sites | at least six people per site |
These figures describe how one specific organization works. I recommend treating them as reference points, not as standards: what matters is setting your own thresholds, writing them down and measuring them.
Sizing the rotation
Section titled “Sizing the rotation”A rotation that is too small mechanically produces fatigue, however good the alerts are. Three questions are enough to check yours.
- How many on-call weeks per person per quarter? With four people on a weekly rotation, each one is on call one week in four, about three weeks per quarter (illustrative example). With eight, it is one week in eight.
- Who is secondary? A designated second level, who knows they may be paged, keeps the primary from being left alone with a long-running incident. The chapter cited above describes this primary and secondary arrangement.
- What happens during holidays and sick leave? If the rotation only holds when everyone is present, it is undersized.
For a team spread across time zones, a follow-the-sun rotation removes night pages. It requires a written handover at every time-zone change.
Handover. I recommend a short, systematic handover at the end of each shift: ongoing incidents, noisy alerts spotted, planned changes, things to watch. Five lines in a shared channel are enough. It also feeds the monthly review.
Measuring alert load
Section titled “Measuring alert load”You cannot reduce what you do not measure. The indicators below come from your alerting tool’s history and from one question asked to the on-call person at the end of the week.
| Indicator | How to compute it | What it reveals |
|---|---|---|
| Pages per shift | alerts that notified a human, per 12-hour period | raw load, to compare with the two-incident benchmark |
| Out-of-hours pages | same count, restricted to nights and weekends | sleep disruption, the most direct burnout factor |
| Share of actionable alerts | alerts that led to a human action, over the total | noise; an alert with no action is a candidate for removal |
| Recurring alerts | alerts fired more than three times in the month with the same cause | what a lasting fix would remove |
| Interrupt time | hours spent on alerts and requests during the shift | what was not spent on engineering |
| Perceived load | a 1 to 5 score given by the person at the end of the shift | drift before it shows in the numbers |
The alert fatigue simulator estimates this load from a few parameters and says whether a typical week stays sustainable.
A thirty-minute monthly review. I recommend reviewing the ten most frequent alerts every month, with three possible decisions per alert: delete it, fix the cause, or turn it into a non-urgent ticket. An alert that comes back three months in a row without a decision is a governance problem, not a technical one.
Only page a human for a symptom
Section titled “Only page a human for a symptom”In my view, the main source of noise is alerting on a cause rather than on a symptom. CPU at 90%, a disk at 80% or a restarted pod do not tell you whether users are suffering. The Monitoring Distributed Systems chapter of the same book recommends alerting on user-visible symptoms and keeping causes for dashboards and diagnosis. It also states that every page should be actionable and warrant an urgent response.
SLOs give a precise criterion for deciding. The Alerting on SLOs chapter of The Site Reliability Workbook describes alerts on the error budget burn rate evaluated over several windows. The principle is:
- fast, sustained budget consumption, which would exhaust it within days, pages a human;
- slow consumption, which would exhaust it before the end of the period without urgency, creates a ticket handled during working hours;
- everything else feeds dashboards, without notification.
The workbook proposes specific thresholds for these windows. The error budget simulator shows how this scheme reduces pages without delaying the detection of a real incident.
flowchart LR
S["Signal"] --> Q1{"Symptom visible<br/>to users?"}
Q1 -->|"no"| D["Dashboard,<br/>diagnosis"]
Q1 -->|"yes"| Q2{"Error budget<br/>burning fast?"}
Q2 -->|"yes"| P["Page on-call"]
Q2 -->|"slowly"| T["Ticket during<br/>working hours"]
Q2 -->|"no"| D
Every alert has an owner and a runbook. I recommend rejecting in review any new paging alert that has neither a link to a runbook (however short) nor an owning team. It is the simplest rule to keep noise from creeping back.
Compensate and protect
Section titled “Compensate and protect”Compensation. In France, article L3121-9 of the Labor Code (Code du travail) defines on-call duty and requires it to be compensated, either financially or with time off; time spent intervening while on call counts as actual working time. The following articles of the same code leave the arrangements to a collective agreement or, failing that, to the employer after consulting employee representatives. The text is available on Légifrance. Other countries have their own rules. I recommend involving HR from the design of the rotation, not after the first complaints.
Recovery. After a night of intervention, I recommend recovery time granted without justification, written into the on-call policy rather than left to each manager’s discretion.
Disconnection. Someone who is not on call does not get paged. Exceptions are named, written down and compensated.
The right to hand back an alert. The on-call person must be able to temporarily silence an obviously noisy alert, provided they flag it in the handover. Without that right, noise settles in because nobody feels entitled to cut it.
Checklist
Section titled “Checklist”- The rotation has enough people to hold during holidays
- A second level is designated and knows it
- Acceptable load thresholds are written down (pages per shift, night pages)
- Alert load is measured every month with the indicators above
- Every paging alert targets a symptom, has an owner and a runbook
- Paging alerts are derived from SLOs where an SLO exists
- Compensation is defined with HR, in line with applicable law
- Recovery after a night intervention is written into the policy
- End-of-shift handover is systematic
- The on-call person can silence a noisy alert by flagging it
Going further
Section titled “Going further”- When on-call ends in an incident, learning goes through the blameless post-mortem.
- Lesson 5 of the CIO path covers on-call together with culture, archetypes and retention.
- For AI services, quality on-call is described in the GenAI method, part VII.
Sources
Section titled “Sources”- Beyer, Jones, Petoff and Murphy (eds.), Site Reliability Engineering, O’Reilly, 2016: chapter 11, Being On-Call and chapter 6, Monitoring Distributed Systems.
- Beyer, Murphy, Rensin, Kawahara and Thorne (eds.), The Site Reliability Workbook, O’Reilly, 2018: chapter 5, Alerting on SLOs and chapter 8, On-Call.
- French Labor Code, article L3121-9, on Légifrance.