Technical
SLOs and alerting
For: engineers and SREs · architects · team managers · business and productPrerequisites: Having read Observability in 10 minutes, or knowing what a metric is.
The previous pages introduced signals and their transport, using the example of a slow payment. One question remains: from what point is this slow payment a problem worth waking someone up for? SLOs give a quantified answer.
The vocabulary and method on this page come from Google’s two books on SRE (Site Reliability Engineering), cited in the sources. SRE is a way of running systems that treats reliability as an engineering problem.
SLI, SLO, SLA in plain words
Section titled “SLI, SLO, SLA in plain words”SLI (service level indicator): a measurement of what the user experiences. It is most often written as a ratio: good events divided by all events. Example: the share of payments that succeed in under 2 seconds.
SLO (service level objective): the target set for that SLI, over a period. Example: 99.9% of payments succeed in under 2 seconds, over a rolling 30 days. It is an internal objective.
SLA (service level agreement): a contractual commitment to a customer, with consequences if it is missed, such as penalties. The SLA is usually less demanding than the SLO. The gap leaves time to react before breaching a contract.
| Term | Question | Example for payment |
|---|---|---|
| SLI | What do I measure? | share of payments succeeding in under 2 s |
| SLO | What level do I aim for? | 99.9% over a rolling 30 days |
| SLA | What did I promise by contract? | 99.5% per month, otherwise a penalty (illustrative) |
A good SLI is measured as close to the user as possible. The success rate seen by the checkout service beats the health status of a server.
The error budget
Section titled “The error budget”Aiming for 99.9% means accepting 0.1% of failures. This tolerated share is called the error budget.
The classic computation goes like this. A 30-day period has 30 × 24 × 60 = 43,200 minutes. 0.1% of 43,200 minutes is 43.2 minutes. A 99.9% SLO over 30 days therefore leaves about 43 minutes of full outage, or the equivalent in partial outage.
| SLO over 30 days | Error budget | Time equivalent |
|---|---|---|
| 99% | 1% | 7 h 12 min |
| 99.5% | 0.5% | 3 h 36 min |
| 99.9% | 0.1% | 43 min |
| 99.95% | 0.05% | 21 min 36 s |
| 99.99% | 0.01% | 4 min 19 s |
These values are computed. Each extra “9” divides the budget by ten.
When the SLI counts requests rather than time, the budget is read in requests. Out of one million payments, 99.9% allows 1,000 failures.
The budget is a decision tool. As long as some remains, the team can ship and take risks. When it is spent, reliability comes first. This rule must be written and agreed in advance: it is the error budget policy.
Alert on symptoms, not causes
Section titled “Alert on symptoms, not causes”A cause is a technical state: a CPU at 90%, a nearly full disk, a restart. A symptom is what the user experiences: failed or slow payments.
The Site Reliability Engineering book recommends alerting on symptoms and keeping causes for diagnosis (Monitoring Distributed Systems). A CPU at 90% that slows no payment does not deserve to wake anyone. Conversely, slow payments deserve an alert, even if every server looks healthy.
In our example, the bank is responding slowly. No machine is saturated. A CPU alert would never have fired. An alert on the payment SLO would have.
Alert on how fast the budget burns
Section titled “Alert on how fast the budget burns”Alerting as soon as an error appears would wake on-call constantly. Waiting for the budget to run out would be too late. The solution is to measure the speed at which the budget is consumed.
This speed is called the burn rate. At 1, the budget runs out exactly at the end of the period. At 2, it runs out in half the time, 15 days over a 30-day period. The higher the burn rate, the more urgent the alert.
The Alerting on SLOs chapter of the Site Reliability Workbook compares several alerting methods. It ends up with multi-window, multi-burn-rate alerts. The principle fits in three ideas.
- Several speeds. Very fast consumption pages on-call. Slow consumption opens a ticket handled during business hours.
- A long window. For each speed, the burn rate is measured over a window long enough to ignore an isolated spike.
- A short window. The burn rate is also checked over a shorter window, about one twelfth of the long one according to the workbook. The alert then stops quickly once the problem is fixed.
For a 99.9% SLO over 30 days, the workbook proposes these parameters as a starting point:
| Action | Long window | Short window | Burn rate | Budget consumed when alerting |
|---|---|---|---|---|
| Page | 1 hour | 5 minutes | 14.4 | 2% |
| Page | 6 hours | 30 minutes | 6 | 5% |
| Ticket | 3 days | 6 hours | 1 | 10% |
To read the first row: at a burn rate of 14.4, one hour consumes 14.4 / 720 = 2% of the 30-day budget (720 hours). At that pace, the whole budget lasts about two days.
flowchart TD
E["SLI error rate"] --> R["Burn rate over<br/>long and short windows"]
R --> Q{"Both windows<br/>above threshold?"}
Q -->|"high burn rate"| P["Page on-call"]
Q -->|"low but sustained burn rate"| T["Ticket"]
Q -->|"no"| D["Dashboard only"]
The error budget simulator applies these rules to an incident you configure yourself. It shows that a sharp outage triggers a page within minutes and that a slow degradation only opens a ticket.
Choosing a first SLO
Section titled “Choosing a first SLO”Here is the approach I recommend for a first SLO. It is deliberately modest.
- Pick a journey that matters. Just one, important to users and to the business. Payment is a good candidate. The home page of an internal tool is less so.
- Write the SLI in plain words. “The share of payments that succeed in under 2 seconds.” If the sentence is not clear to the business, the SLI is not either.
- Measure before aiming. Look at the actual reliability over the last few weeks. A target picked at random will be either unreachable or pointless.
- Set a target slightly below what you observe. If the service holds 99.95%, aiming for 99.9% leaves a margin. Aiming for 100% makes no sense: no budget, no possible release.
- Write the budget policy. Who decides to slow down releases when the budget is spent? Before the incident, not during it.
- Wire burn rate alerts to this SLO, then remove the cause-based alerts that duplicate them.
- Review after a quarter. The first SLO is rarely the right one. That is normal.
Going further
Section titled “Going further”- Error budget simulator: set an SLO, simulate an incident, watch the alerts fire.
- Sustainable on-call: what SLO-based alerts change for the people on call.
- Alert fatigue simulator: estimate the load of an on-call week.
- Quality SLO: apply the same logic to the quality of an AI service’s answers.
- Observability in 10 minutes and OpenTelemetry end to end: the signals SLIs are built on.
- The glossary defines SLI, SLO, error budget and burn rate.
Sources
Section titled “Sources”- Beyer, Jones, Petoff and Murphy (eds.), Site Reliability Engineering, O’Reilly, 2016: chapter 4, Service Level Objectives and chapter 6, Monitoring Distributed Systems.
- Beyer, Murphy, Rensin, Kawahara and Thorne (eds.), The Site Reliability Workbook, O’Reilly, 2018: chapter 2, Implementing SLOs and chapter 5, Alerting on SLOs (table of recommended parameters for a 99.9% SLO over 30 days).
- The budgets in the “SLO over 30 days” table are computed from a 43,200-minute period.