People
People dimension
A well-built observability platform can end up looked at by nobody, and an on-call rotation paged for nothing too often ends up no longer reacting to alerts. The people dimension is about people, their reflexes and their limits.
The chapters follow the learning loop of an incident.
flowchart TB A["Actionable<br/>alert"] --> R["Response<br/>and diagnosis"] --> RS["Recovery"] --> PM["Blameless<br/>post-mortem"] PM --> AC["Actions,<br/>not culprits"] AC -->|"thresholds, dashboards,<br/>runbooks"| A PM -.->|"real cost of the incident"| MET["Business lines<br/>concerned"]
Published and upcoming chapters
Section titled “Published and upcoming chapters”| Chapter | What you learn to do | Status |
|---|---|---|
| Alert fatigue | measure noise, remove non-actionable alerts | planned |
| Sustainable on-call | size rotations, protect sleep, tool escalation | published; the leadership view in the CIO path, lesson 5; quality on-call for AI in the GenAI method |
| Blameless post-mortems | run an incident review that produces actions, not culprits | published; four cultural rules in the CIO path, lesson 5 |
| Troubleshooting under pressure | know the cognitive biases at play and design dashboards that limit them | planned |
| Roles and skills | define roles (SRE, observability engineer, platform owner) and their career paths | archetypes and career paths in the CIO path, lesson 5 |
| Driving adoption | bring development teams on board, measure real usage | ADKAR applied in the CIO path, lesson 6; dedicated chapter planned |
Measuring it
Section titled “Measuring it”The alert fatigue simulator estimates the real load of an on-call week and says whether it is sustainable. The error budget simulator shows how multi-window alerts wake people up less often for nothing.