People
Blameless post-mortems
For: engineers and SREs · team managers · business and product · executives and CIOsPrerequisites: Basic notions of incident management.
An incident costs you anyway. The only question is whether the organization gets something out of it. The blameless post-mortem is the ritual that turns an incident into learning: it looks for what made the error possible rather than for the person who made it, then produces a small number of actions followed through to the end.
Lesson 5 of the CIO path sets out four cultural rules. This article is the facilitation guide: when to trigger, how to run the session, which template to use, how to track actions and what to measure.
Why “blameless”
Section titled “Why “blameless””The argument is practical before it is moral. If someone risks a sanction by describing what they did, they will say less, or say it differently. The analysis then stops at “human error” and the condition that made the error possible stays in place for the next person. John Allspaw put this idea into words for Etsy in a text that became a reference (Blameless PostMortems and a Just Culture), and the Post-mortem Culture chapter of the Site Reliability Engineering book builds on it.
Blameless does not mean without accountability. Actions have owners and dates. What disappears is the search for a culprit.
When to trigger a post-mortem
Section titled “When to trigger a post-mortem”A post-mortem costs several people a few hours of work. You therefore need written criteria, known before the incident, so it depends neither on mood nor on perceived severity. The Google chapter cited above proposes triggers that I use as a baseline:
- user-visible downtime or degradation beyond a threshold set in advance;
- data loss of any kind;
- a manual on-call intervention to restore service (rollback, traffic rerouting);
- a resolution time above a threshold;
- an incident discovered by something other than monitoring, which signals an observability gap.
The same chapter adds that any stakeholder may request a post-mortem, even outside the criteria. I add one recommendation: an incident that consumes a significant share of an SLO’s error budget, a share you set, triggers a post-mortem even if nobody complained.
The process
Section titled “The process”| Step | When | Who | Output |
|---|---|---|---|
| Assignment | when the incident is closed | the incident lead | an author and a facilitator, ideally not involved in the response |
| Collection | in the following days | the author | timestamped timeline, dashboard screenshots, incident channel excerpts |
| Draft | before the session | the author, reviewed by responders | the document following the template, action section empty |
| Session | about an hour | responders, owning team, affected business line | causes and factors validated, actions decided and assigned |
| Publication | after the session | the author | document shared internally, actions entered in the backlog |
| Follow-up | until closure | action owners, reviewed monthly | actions closed or explicitly dropped |
I recommend setting a target delay between incident closure and the session, short enough for memories to be fresh. The right delay depends on your pace; what matters is that it is written down and measured.
The facilitator’s role. They protect the rule. They rephrase “they should have” into “what made this action reasonable at that moment?”. They bring the discussion back to conditions (tooling, procedure, missing signal, time pressure) when it drifts towards people. They make sure the most junior participants speak first.
Timeline first. I recommend spending the first third of the session validating the timeline. Disagreements about causes often come from unspoken disagreements about facts.
The template
Section titled “The template”A short template gets used; a long one gets bypassed. This one fits on two pages.
# Post-mortem: <factual title, no person's name>
Status: draft | reviewed | publishedIncident date: YYYY-MM-DD Duration: from HH:MM to HH:MMAuthor: Facilitator:Services affected: SLOs concerned:
## SummaryThree sentences: what happened, the impact, how service was restored.
## Impact- Users or customers affected, over which period- Error budget consumed- Business impact established with the affected business line (orders, delays, reputation)
## TimelineHH:MM event (source: alert, channel, dashboard)Detection, first diagnosis, decisions, recovery.
## DetectionHow the incident was detected. How long after it started.Did monitoring see it? If not, which signal would have been enough?
## Causes and contributing factorsWhat triggered the incident. What made it possible.What made it worse or longer.
## What went well## Where we got lucky
## Actions| Action | Type (prevent, detect, mitigate) | Owner | Due date | Done when |
## LessonsWhat other teams should know.Two sections deserve particular attention. Detection links the post-mortem to observability: an incident detected by a customer rather than an alert is an instrumentation action in itself. Where we got lucky surfaces the risks that did not materialize this time, often the most instructive ones.
Actions, where everything is decided
Section titled “Actions, where everything is decided”A post-mortem whose actions are never carried out is worse than no post-mortem: it teaches teams that the ritual is useless.
I recommend four rules:
- Few, chosen actions. Three completed actions beat twelve pending ones. Ideas not retained go into a “leads” section with no owner.
- One action, one owner, one date, one completion criterion. “Improve monitoring” is not an action. “Add an alert on the payment service error rate, tied to its SLO” is one.
- Actions live in the team’s backlog, not in the document. They are prioritized like the rest of the work, with a label that makes them easy to find.
- A monthly review of open actions, short, that closes, chases or explicitly drops them. A deliberate, justified drop is better than an action left to rot.
The Post-mortem Culture chapter of The Site Reliability Workbook compares a weak post-mortem with an improved version. I recommend reading it to anyone who will facilitate sessions.
What to measure
Section titled “What to measure”These indicators tell you whether the ritual works. They are not for ranking teams: a high number of post-mortems may signal a healthy reporting culture rather than a fragile system.
| Indicator | What it reveals |
|---|---|
| Share of eligible incidents with a published post-mortem | whether the trigger criteria are applied |
| Delay between incident closure and publication | whether the ritual is kept or postponed |
| Share of actions closed by their due date | whether the post-mortem actually changes the system |
| Age of open actions | accumulated learning debt |
| Recurring incidents (same cause as an incident already analyzed) | whether actions addressed the right cause |
| Share of incidents detected by monitoring rather than by a user | what post-mortems bring to observability |
| Reads or shares of published post-mortems | whether learning reaches beyond the team concerned |
Recurring incidents and the share detected by monitoring can feed the executive dashboard: they turn learning into indicators a leadership committee understands.
Common pitfalls
Section titled “Common pitfalls”- The post-mortem as a trial. A manager asking “who did this?” in the session cancels the rule for a long time. The facilitator must be able to stop it.
- The single root cause. Incidents in distributed systems rarely have a single cause. I recommend talking about contributing factors.
- Human error as the conclusion. It is a starting point: what made the error easy and its detection hard?
- The document nobody reads. Broad internal sharing and a three-sentence summary do more than an exhaustive document filed away in a folder.
Checklist
Section titled “Checklist”- Trigger criteria are written down and known before the incident
- The facilitator did not take part in the incident response
- The timeline is validated before any discussion of causes
- Neither the title nor the document names a culprit
- The Detection section says whether monitoring saw the incident
- Every action has an owner, a due date and a completion criterion
- Actions are in the team’s backlog
- Open actions are reviewed every month
- The affected business line is invited when the incident hit it
Going further
Section titled “Going further”- Post-mortems depend on sustainable on-call: an exhausted team no longer has the energy to analyze.
- The MTTR business case links detection and recovery time to the cost of an incident.
- CIO path, lesson 5 offers a replayed post-mortem exercise to measure cultural maturity.
Sources
Section titled “Sources”- Beyer, Jones, Petoff and Murphy (eds.), Site Reliability Engineering, O’Reilly, 2016: chapter 15, Post-mortem Culture: Learning from Failure.
- Beyer, Murphy, Rensin, Kawahara and Thorne (eds.), The Site Reliability Workbook, O’Reilly, 2018: chapter 10, Post-mortem Culture: Learning from Failure.
- John Allspaw, Blameless PostMortems and a Just Culture, Code as Craft, Etsy, 2012.
- Public examples of processes and templates: PagerDuty’s post-mortem guide and Atlassian’s incident management handbook.