Skip to content

People

People dimension

A well-built observability platform can end up looked at by nobody, and an on-call rotation paged for nothing too often ends up no longer reacting to alerts. The people dimension is about people, their reflexes and their limits.

The chapters follow the learning loop of an incident.

flowchart TB
  A["Actionable<br/>alert"] --> R["Response<br/>and diagnosis"] --> RS["Recovery"] --> PM["Blameless<br/>post-mortem"]
  PM --> AC["Actions,<br/>not culprits"]
  AC -->|"thresholds, dashboards,<br/>runbooks"| A
  PM -.->|"real cost of the incident"| MET["Business lines<br/>concerned"]
ChapterWhat you learn to doStatus
Alert fatiguemeasure noise, remove non-actionable alertsplanned
Sustainable on-callsize rotations, protect sleep, tool escalationpublished; the leadership view in the CIO path, lesson 5; quality on-call for AI in the GenAI method
Blameless post-mortemsrun an incident review that produces actions, not culpritspublished; four cultural rules in the CIO path, lesson 5
Troubleshooting under pressureknow the cognitive biases at play and design dashboards that limit themplanned
Roles and skillsdefine roles (SRE, observability engineer, platform owner) and their career pathsarchetypes and career paths in the CIO path, lesson 5
Driving adoptionbring development teams on board, measure real usageADKAR applied in the CIO path, lesson 6; dedicated chapter planned

The alert fatigue simulator estimates the real load of an on-call week and says whether it is sustainable. The error budget simulator shows how multi-window alerts wake people up less often for nothing.