Skip to content

PeoplePractitioner

Sustainable on-call

For: engineers and SREs · team managersPrerequisites: Basic notions of alerting and SLOs.

Reading mode

On-call is not sustainable because the team is brave. It is sustainable because four things hold at the same time: a large enough rotation, few alerts that are all actionable, recognized compensation and regular measurement of the real load. When one is missing, the others eventually give way.

This article gathers the published benchmarks, a measurement method and the rules I recommend. It complements lesson 5 of the CIO path, which looks at the topic from the leadership side.

Google’s Site Reliability Engineering book has a chapter on on-call (Being On-Call). It gives orders of magnitude that are widely quoted:

BenchmarkValue proposed by the book
Share of an engineer’s time spent on callat most 25%
Time spent on engineering (not operations)at least 50%
Incidents per 12-hour on-call shiftat most two
Rotation size, single siteat least eight people
Rotation size, two sitesat least six people per site

These figures describe how one specific organization works. I recommend treating them as reference points, not as standards: what matters is setting your own thresholds, writing them down and measuring them.

A rotation that is too small mechanically produces fatigue, however good the alerts are. Three questions are enough to check yours.

  1. How many on-call weeks per person per quarter? With four people on a weekly rotation, each one is on call one week in four, about three weeks per quarter (illustrative example). With eight, it is one week in eight.
  2. Who is secondary? A designated second level, who knows they may be paged, keeps the primary from being left alone with a long-running incident. The chapter cited above describes this primary and secondary arrangement.
  3. What happens during holidays and sick leave? If the rotation only holds when everyone is present, it is undersized.

For a team spread across time zones, a follow-the-sun rotation removes night pages. It requires a written handover at every time-zone change.

Handover. I recommend a short, systematic handover at the end of each shift: ongoing incidents, noisy alerts spotted, planned changes, things to watch. Five lines in a shared channel are enough. It also feeds the monthly review.

You cannot reduce what you do not measure. The indicators below come from your alerting tool’s history and from one question asked to the on-call person at the end of the week.

IndicatorHow to compute itWhat it reveals
Pages per shiftalerts that notified a human, per 12-hour periodraw load, to compare with the two-incident benchmark
Out-of-hours pagessame count, restricted to nights and weekendssleep disruption, the most direct burnout factor
Share of actionable alertsalerts that led to a human action, over the totalnoise; an alert with no action is a candidate for removal
Recurring alertsalerts fired more than three times in the month with the same causewhat a lasting fix would remove
Interrupt timehours spent on alerts and requests during the shiftwhat was not spent on engineering
Perceived loada 1 to 5 score given by the person at the end of the shiftdrift before it shows in the numbers

The alert fatigue simulator estimates this load from a few parameters and says whether a typical week stays sustainable.

A thirty-minute monthly review. I recommend reviewing the ten most frequent alerts every month, with three possible decisions per alert: delete it, fix the cause, or turn it into a non-urgent ticket. An alert that comes back three months in a row without a decision is a governance problem, not a technical one.

In my view, the main source of noise is alerting on a cause rather than on a symptom. CPU at 90%, a disk at 80% or a restarted pod do not tell you whether users are suffering. The Monitoring Distributed Systems chapter of the same book recommends alerting on user-visible symptoms and keeping causes for dashboards and diagnosis. It also states that every page should be actionable and warrant an urgent response.

SLOs give a precise criterion for deciding. The Alerting on SLOs chapter of The Site Reliability Workbook describes alerts on the error budget burn rate evaluated over several windows. The principle is:

  • fast, sustained budget consumption, which would exhaust it within days, pages a human;
  • slow consumption, which would exhaust it before the end of the period without urgency, creates a ticket handled during working hours;
  • everything else feeds dashboards, without notification.

The workbook proposes specific thresholds for these windows. The error budget simulator shows how this scheme reduces pages without delaying the detection of a real incident.

flowchart LR
  S["Signal"] --> Q1{"Symptom visible<br/>to users?"}
  Q1 -->|"no"| D["Dashboard,<br/>diagnosis"]
  Q1 -->|"yes"| Q2{"Error budget<br/>burning fast?"}
  Q2 -->|"yes"| P["Page on-call"]
  Q2 -->|"slowly"| T["Ticket during<br/>working hours"]
  Q2 -->|"no"| D

Every alert has an owner and a runbook. I recommend rejecting in review any new paging alert that has neither a link to a runbook (however short) nor an owning team. It is the simplest rule to keep noise from creeping back.

Compensation. In France, article L3121-9 of the Labor Code (Code du travail) defines on-call duty and requires it to be compensated, either financially or with time off; time spent intervening while on call counts as actual working time. The following articles of the same code leave the arrangements to a collective agreement or, failing that, to the employer after consulting employee representatives. The text is available on Légifrance. Other countries have their own rules. I recommend involving HR from the design of the rotation, not after the first complaints.

Recovery. After a night of intervention, I recommend recovery time granted without justification, written into the on-call policy rather than left to each manager’s discretion.

Disconnection. Someone who is not on call does not get paged. Exceptions are named, written down and compensated.

The right to hand back an alert. The on-call person must be able to temporarily silence an obviously noisy alert, provided they flag it in the handover. Without that right, noise settles in because nobody feels entitled to cut it.

  • The rotation has enough people to hold during holidays
  • A second level is designated and knows it
  • Acceptable load thresholds are written down (pages per shift, night pages)
  • Alert load is measured every month with the indicators above
  • Every paging alert targets a symptom, has an owner and a runbook
  • Paging alerts are derived from SLOs where an SLO exists
  • Compensation is defined with HR, in line with applicable law
  • Recovery after a night intervention is written into the policy
  • End-of-shift handover is systematic
  • The on-call person can silence a noisy alert by flagging it