Skip to content

PeoplePractitioner

Blameless post-mortems

For: engineers and SREs · team managers · business and product · executives and CIOsPrerequisites: Basic notions of incident management.

Reading mode

An incident costs you anyway. The only question is whether the organization gets something out of it. The blameless post-mortem is the ritual that turns an incident into learning: it looks for what made the error possible rather than for the person who made it, then produces a small number of actions followed through to the end.

Lesson 5 of the CIO path sets out four cultural rules. This article is the facilitation guide: when to trigger, how to run the session, which template to use, how to track actions and what to measure.

The argument is practical before it is moral. If someone risks a sanction by describing what they did, they will say less, or say it differently. The analysis then stops at “human error” and the condition that made the error possible stays in place for the next person. John Allspaw put this idea into words for Etsy in a text that became a reference (Blameless PostMortems and a Just Culture), and the Post-mortem Culture chapter of the Site Reliability Engineering book builds on it.

Blameless does not mean without accountability. Actions have owners and dates. What disappears is the search for a culprit.

A post-mortem costs several people a few hours of work. You therefore need written criteria, known before the incident, so it depends neither on mood nor on perceived severity. The Google chapter cited above proposes triggers that I use as a baseline:

  • user-visible downtime or degradation beyond a threshold set in advance;
  • data loss of any kind;
  • a manual on-call intervention to restore service (rollback, traffic rerouting);
  • a resolution time above a threshold;
  • an incident discovered by something other than monitoring, which signals an observability gap.

The same chapter adds that any stakeholder may request a post-mortem, even outside the criteria. I add one recommendation: an incident that consumes a significant share of an SLO’s error budget, a share you set, triggers a post-mortem even if nobody complained.

StepWhenWhoOutput
Assignmentwhen the incident is closedthe incident leadan author and a facilitator, ideally not involved in the response
Collectionin the following daysthe authortimestamped timeline, dashboard screenshots, incident channel excerpts
Draftbefore the sessionthe author, reviewed by respondersthe document following the template, action section empty
Sessionabout an hourresponders, owning team, affected business linecauses and factors validated, actions decided and assigned
Publicationafter the sessionthe authordocument shared internally, actions entered in the backlog
Follow-upuntil closureaction owners, reviewed monthlyactions closed or explicitly dropped

I recommend setting a target delay between incident closure and the session, short enough for memories to be fresh. The right delay depends on your pace; what matters is that it is written down and measured.

The facilitator’s role. They protect the rule. They rephrase “they should have” into “what made this action reasonable at that moment?”. They bring the discussion back to conditions (tooling, procedure, missing signal, time pressure) when it drifts towards people. They make sure the most junior participants speak first.

Timeline first. I recommend spending the first third of the session validating the timeline. Disagreements about causes often come from unspoken disagreements about facts.

A short template gets used; a long one gets bypassed. This one fits on two pages.

# Post-mortem: <factual title, no person's name>
Status: draft | reviewed | published
Incident date: YYYY-MM-DD Duration: from HH:MM to HH:MM
Author: Facilitator:
Services affected: SLOs concerned:
## Summary
Three sentences: what happened, the impact, how service was restored.
## Impact
- Users or customers affected, over which period
- Error budget consumed
- Business impact established with the affected business line (orders, delays, reputation)
## Timeline
HH:MM event (source: alert, channel, dashboard)
Detection, first diagnosis, decisions, recovery.
## Detection
How the incident was detected. How long after it started.
Did monitoring see it? If not, which signal would have been enough?
## Causes and contributing factors
What triggered the incident. What made it possible.
What made it worse or longer.
## What went well
## Where we got lucky
## Actions
| Action | Type (prevent, detect, mitigate) | Owner | Due date | Done when |
## Lessons
What other teams should know.

Two sections deserve particular attention. Detection links the post-mortem to observability: an incident detected by a customer rather than an alert is an instrumentation action in itself. Where we got lucky surfaces the risks that did not materialize this time, often the most instructive ones.

A post-mortem whose actions are never carried out is worse than no post-mortem: it teaches teams that the ritual is useless.

I recommend four rules:

  1. Few, chosen actions. Three completed actions beat twelve pending ones. Ideas not retained go into a “leads” section with no owner.
  2. One action, one owner, one date, one completion criterion. “Improve monitoring” is not an action. “Add an alert on the payment service error rate, tied to its SLO” is one.
  3. Actions live in the team’s backlog, not in the document. They are prioritized like the rest of the work, with a label that makes them easy to find.
  4. A monthly review of open actions, short, that closes, chases or explicitly drops them. A deliberate, justified drop is better than an action left to rot.

The Post-mortem Culture chapter of The Site Reliability Workbook compares a weak post-mortem with an improved version. I recommend reading it to anyone who will facilitate sessions.

These indicators tell you whether the ritual works. They are not for ranking teams: a high number of post-mortems may signal a healthy reporting culture rather than a fragile system.

IndicatorWhat it reveals
Share of eligible incidents with a published post-mortemwhether the trigger criteria are applied
Delay between incident closure and publicationwhether the ritual is kept or postponed
Share of actions closed by their due datewhether the post-mortem actually changes the system
Age of open actionsaccumulated learning debt
Recurring incidents (same cause as an incident already analyzed)whether actions addressed the right cause
Share of incidents detected by monitoring rather than by a userwhat post-mortems bring to observability
Reads or shares of published post-mortemswhether learning reaches beyond the team concerned

Recurring incidents and the share detected by monitoring can feed the executive dashboard: they turn learning into indicators a leadership committee understands.

  • The post-mortem as a trial. A manager asking “who did this?” in the session cancels the rule for a long time. The facilitator must be able to stop it.
  • The single root cause. Incidents in distributed systems rarely have a single cause. I recommend talking about contributing factors.
  • Human error as the conclusion. It is a starting point: what made the error easy and its detection hard?
  • The document nobody reads. Broad internal sharing and a three-sentence summary do more than an exhaustive document filed away in a folder.
  • Trigger criteria are written down and known before the incident
  • The facilitator did not take part in the incident response
  • The timeline is validated before any discussion of causes
  • Neither the title nor the document names a culprit
  • The Detection section says whether monitoring saw the incident
  • Every action has an owner, a due date and a completion criterion
  • Actions are in the team’s backlog
  • Open actions are reviewed every month
  • The affected business line is invited when the incident hit it