Key takeaways
- →Blameless does not mean consequence-free - it means the review's output is an explanation of the system, not a judgement about a person.
- →Agree the timeline as pure fact before anyone offers a cause, because the first explanation spoken out loud anchors every explanation after it.
- →Replace every question starting with 'who' or 'why did you' with a question about what the system showed and hid at that moment.
- →Collect written contributions silently before the discussion so the on-call engineer is not the only person supplying the narrative.
- →Leave with three actions that have named owners and dates, not eleven that have neither - unowned actions are how reviews become theatre.
- →Track whether people report near-misses between incidents; that is the real measure of whether your reviews feel safe.
A post-incident review is blameless when the group can explain why a decision looked correct to the person making it at the time, given the information they had and the pressure they were under. If the meeting ends with a name and a promise that someone will be more careful, it was not blameless, and it has not made the next incident less likely.
That distinction matters for a practical reason rather than a cultural one. The detail you need to fix a system - the alert that was ignored because it fires forty times a week, the runbook step everyone quietly skips, the dashboard that was open on the wrong environment - only surfaces when the person who holds it does not expect to be punished for saying it out loud. Blame does not make people lie; it makes them go quiet and factual, and factual is not enough.
This guide covers what blameless actually means and what it does not, when to hold the review, a sixty-minute agenda, how to build a timeline without inviting explanations too early, the question set that replaces 'who', what to do when blame enters the room anyway, and how to turn the output into a small number of actions that get done.
What blameless actually means - and what it does not
Blameless means the review produces an explanation of the system rather than a verdict on a person. The working assumption is that everyone involved acted sensibly with what they could see, so any behaviour that looks obviously wrong afterwards is a signal that the system made the wrong thing easy, fast or invisible. That assumption is not charity - it is the only assumption that gets you the information you need.
It does not mean nobody is accountable. The team is accountable for the fix, the owner of each action is accountable for delivering it, and the organisation is accountable for the conditions that made the failure likely. What is off the table is using an hour of collective attention to work out whose fault it was, because that answer has never once prevented a recurrence.
It also does not mean that genuine performance or conduct problems disappear. If someone deliberately bypassed a control or repeatedly ignored an agreed process, that is a conversation between a manager and an individual, held separately and privately. Say this out loud at the start of the review. Naming the boundary is what makes the room believe the boundary exists.
The hardest part is hindsight. Once you know the outcome, the path to it looks like a lit corridor, and every branch not taken looks like negligence. At the time it was one of six plausible readings of an ambiguous graph at 02:40. A good facilitator keeps saying so, in those words, whenever the room starts talking as though the answer was available.
Rewrite the questions before you rewrite the process
Most reviews turn adversarial through wording rather than intent. The left column is what people say when they are genuinely trying to understand; the right column is the same curiosity aimed at the system instead of the person. Print the right-hand column and keep it in front of you while you facilitate.
The pattern is simple enough to apply live. Any question containing 'you' as the subject of a failure becomes a question about what was on the screen, in the runbook, in the alert or in the calendar at that moment. It takes about three sessions before the team starts rewriting each other's questions without you.
| Question people ask | What it produces | The version that works |
|---|---|---|
| Why did you deploy on a Friday? | Defensiveness, and the actual reason stays unsaid | What made Friday afternoon look like a reasonable time to ship this? |
| Who approved this change? | A name for the report and nothing else | What did the approval step check, and what could it not see? |
| Why did nobody notice for forty minutes? | Silence from whoever was on call | What signals existed at minute five, and where were they visible? |
| How did this get through code review? | Reviewers who stop volunteering for risky changes | What would a reviewer have needed in front of them to catch this? |
| Was the runbook followed? | Compliance theatre during the next incident | Where did the runbook and the live system disagree? |
| Whose mistake was it? | A culprit, an unchanged system and a repeat in six weeks | What made this mistake easy to make and hard to catch? |
Hold the review within five working days
Two to five working days after resolution is the window for most incidents. Any sooner and people are still exhausted, still cleaning up, and still close enough to the adrenaline that the conversation runs hot. Any later and the detail decays: log retention rolls over, chat threads scroll away, and - more damaging - everyone has privately settled on a story that the meeting will now only confirm.
Do one thing on day one regardless: capture the raw timeline while it is cheap. A single shared document, appended to by whoever was involved, with timestamps and no interpretation. Ten minutes of copy-paste from the incident channel on the day saves you twenty-five minutes of collective archaeology in the review itself.
Pick a facilitator who did not work the incident. Someone who spent the night fixing it cannot simultaneously run a neutral conversation about it - they will defend decisions they made at 03:00 without noticing they are doing it. A facilitator from a neighbouring team is usually better than a senior person from the same one, because seniority in the room quietly narrows what gets said.
Keep the invite list to the people who were involved, the people who own the affected system, and one or two who will genuinely learn from it. Eight to twelve is workable. Above about fifteen the review turns into a briefing, contributions concentrate in three voices, and the quiet engineer who noticed the odd metric at minute two never mentions it.
A sixty-minute agenda for a single-service incident
Sixty minutes is right for a single-service incident with a clear start and end. A multi-team outage that ran for hours needs ninety minutes and a pre-read, and something that touched customers, money or data needs a second session a week later once the technical picture has settled.
Notice how little of the hour is open discussion. Fifteen minutes of it is a facilitator reading a timeline, fifteen is people typing in silence, and only fifteen is the free-form conversation most teams currently spend the entire hour on. That ratio is the single biggest change you can make.
0:00 - 0:05 Purpose and boundary
State the outcome, the finish time and the one rule. Say explicitly that performance conversations happen elsewhere and that nothing here goes to anyone's manager as evidence.
0:05 - 0:20 Timeline, facts only
Walk the timestamped sequence out loud. Corrections and additions welcome; explanations parked. Mark each point with what was known then, not what is known now.
0:20 - 0:35 Silent written contributions
Everyone types answers to three prompts at once - what made this hard to see, what made it hard to fix, what made it easy to cause. No talking for the first five minutes.
0:35 - 0:50 Discuss the top themes
Rank or vote on the written contributions, then discuss only the top three or four. This is the only part of the hour that is open conversation, and it is deliberately short.
0:50 - 0:58 Actions, owners, dates
Three actions maximum. Each one has a named human and a date said out loud in the room. Anything you cannot assign an owner to goes on the backlog, not the action list.
0:58 - 1:00 Close with a pulse
One anonymous question: was anything left unsaid today? A single word or a five-point scale is enough. It takes twenty seconds and it is the cheapest safety check you have.
Build the timeline before anyone explains anything
The order here is the whole trick. Timeline first, explanations second, and a hard boundary between them - because the first cause spoken out loud sets the frame for everything after it. If the first sentence in the room is 'this is really a caching problem', the next forty minutes will be about caching, whether or not caching was the interesting part.
That anchoring effect is not a failure of intelligence; it is how groups converge. The counter is procedural rather than motivational: keep the causal conversation closed until everyone has agreed on what happened and when, and collect the candidate causes in writing so they all arrive at once instead of in order of who speaks fastest.
Assemble the raw events in advance
Deploy times, alert times, first human acknowledgement, first customer report, each mitigation attempt and the moment recovery was confirmed. Pull them from logs and chat rather than memory, and circulate the document before the meeting so people arrive arguing about facts rather than discovering them.
Read it out loud in order
Reading the sequence aloud, slowly, does something that skimming a document does not: it exposes the gaps. Twelve minutes between the alert and the first acknowledgement sounds very different when you say it than when it sits in a table nobody has looked at.
Annotate what each person knew at each point
At every significant moment, ask what was visible to whoever was acting. Not what was true - what was visible. This is where you discover that the dashboard everyone assumed was being watched had been broken since a migration three weeks earlier.
Mark the ambiguous moments explicitly
Any point where a reasonable person could have read the situation two ways gets a marker. These markers, not the eventual root cause, are where your most useful actions come from, because ambiguity is fixable with better signals and clearer defaults.
Park every explanation until the timeline is agreed
When someone offers a cause during the timeline phase - and they will, in the first four minutes - write it down visibly and say you will come back to it. Parking it publicly is the move; refusing it without recording it makes people stop offering things.
Ask what would have to be true for it to be worse
Before you leave the timeline, ask what would have made this incident three times larger. The answer usually names your next investment more accurately than the review of the incident that actually happened.
The question that carries the whole review
If you only remember one prompt, make it this: what made this look like the right thing to do at the time? Ask it every time the room reaches a decision that looks indefensible in hindsight, and ask it neutrally, as though the answer is genuinely interesting - because it is.
The answers are consistently the useful part of the meeting. The engineer restarted the wrong service because the two are named one character apart in the console. The alert was silenced because the same alert had cried wolf eleven times that fortnight. The rollback was not attempted because nobody present had ever seen it work and the runbook had not been touched in a year. None of those are careless people; all of them are fixable systems.
There is a second prompt worth keeping in reserve for the quiet moments: what nearly stopped this from happening? Near-misses and lucky catches are as instructive as the failure itself, and they are the part of the story people rarely volunteer because nothing went wrong in that branch.
The purpose of a post-incident review is not to establish who was at the keyboard. It is to establish what the system made obvious to the person at the keyboard, and what it kept hidden.
Collect contributions in writing before you discuss
The default review format - facilitator asks an open question, room responds verbally - collects from whoever is most senior, most confident and most involved. In a post-incident review that is almost always the on-call engineer who has already replayed the night in their head fifty times, and their narrative becomes the group's narrative within the first three minutes.
Silent, simultaneous, written generation fixes this, and it is the oldest trick in structured facilitation. Give everyone the same three prompts, ten to fifteen minutes with no talking, then surface every contribution at once. You get input from the person who was on the periphery and noticed something odd, from the newer engineer who found the runbook incomprehensible, and from whoever disagrees with the emerging story but would not have interrupted to say so.
Make that phase anonymous when the incident touched anything sensitive - a customer-facing outage, a data issue, anything where somebody is quietly worried about their job. Session Flo runs the whole sequence from one room code: an anonymous open-text collection for the three prompts, a ranking round so the group picks what to discuss rather than the loudest voice picking it, and a closing pulse poll on whether anything went unsaid. Everything lands in a recap you can attach to the incident record.
Then hold to the ranking. If the room voted the alerting noise to the top and the deploy process third, spend the discussion on alerting noise. Overriding the ranking because you personally think the deploy process matters more teaches everyone that the vote was decorative, and next time they will not bother.
When blame walks into the room anyway
Top Tips
- Redirect in the moment rather than after the meeting. A single calm sentence - 'let us look at what the console showed rather than who clicked' - costs four seconds and resets the norm for the rest of the hour. Saying nothing is a decision the room will read correctly.
- Watch for self-blame, which is more common and more corrosive than blaming others. When someone says 'that one is on me, I should have checked', accept it once and immediately turn it outward: what would have had to exist for checking to be automatic rather than remembered?
- Treat 'human error' as the beginning of the analysis rather than the end of it. It is a category, not a cause. Every time it appears in a draft write-up, send it back with one question: what made the error easy to make and slow to detect?
- Handle the senior person who arrives looking for accountability by giving them a job. Ask them to own one of the three actions. A director with an action on their name behaves very differently from a director spectating with folded arms.
- Do not let the review become a design meeting. When a fix debate runs past three minutes, note the decision that needs making, name an owner, and move on. Reviews that turn into architecture arguments never get to the actions.
- If the incident involved another team, invite them and brief them beforehand on the format. A team walking into an unfamiliar blameless review with no warning will assume it is a tribunal and behave accordingly, which is exactly the dynamic you are trying to avoid.
Running the review remotely or hybrid
Remote reviews are actually easier to keep blameless than in-person ones, provided you use the medium properly. Written contribution is the natural format on a call, the timeline can sit on a shared screen that everyone reads at the same pace, and anonymity is trivial in a way that it is not around a table.
Hybrid is where reviews go wrong. If four people are in a room and five are on a call, the in-room conversation moves faster, the side comments never reach the remote participants, and the person who worked the incident from home ends up hearing conclusions rather than shaping them. The fix is mechanical: every contribution goes through the device, including from the people sitting together, and the timeline is driven by someone remote.
Keep cameras optional and say so. People discussing a night they handled badly do not need to also manage their face for an hour, and the participation cost of forcing cameras is higher here than in almost any other kind of session. What you need is their typing, not their expression.
Numbers worth holding the review to
These are structural guidelines rather than research findings - they are the shape that keeps a review honest, and they are worth defending when someone suggests inviting thirty people or leaving the write-up until next sprint.
The action count is the one people push back on hardest. Eleven actions feels thorough and produces nothing; three actions with names and dates changes something by the end of the month. If the review genuinely surfaced eleven things worth doing, that is a planning conversation, not an outcome of the hour.
Turn findings into actions that actually get done
An action is only real if it has a person, a date and a place it lives outside the document. 'Improve monitoring' is not an action; 'Ana adds a synthetic check on the checkout path, in the sprint starting the 24th, tracked as ENG-4471' is. Say the owner's name out loud in the room and let them accept or push back on the date while everyone is still present.
Sort the candidates into three buckets before you assign anything. Reduce likelihood, reduce detection time, reduce impact. Teams overwhelmingly pick the first bucket, which is the hardest and slowest, when the fastest wins are usually in the second - an alert that fires ten minutes earlier converts a forty-minute outage into a five-minute one without preventing anything.
Then close the loop somewhere visible. Review open incident actions at the start of the next one, or in a standing slot in the team's planning meeting. Nothing kills the credibility of blameless reviews faster than a team that runs them faithfully and can point to nine months of actions that were never delivered - at that point people are attending a ritual, and they will contribute the way you contribute to a ritual.
What the write-up has to contain
Write for someone two teams away who has the same failure mode and does not know it yet. That reader needs the customer impact, the conditions and the actions; they do not need forty lines of stack trace, and they certainly do not need to know who was on call.
Use roles rather than names throughout - 'the on-call engineer', 'the deploying team', 'the reviewer'. It reads slightly stiffer and it removes the single most common reason people sanitise their contributions: knowing that a searchable document with their name in it will outlive the incident by years.
The last item is the one teams skip. An internal review that only its own team reads is worth a fraction of one that a neighbouring team can find, because the same broken defaults are almost always sitting in three other services.
- ✓A plain-language summary of what customers experienced, and for how long
- ✓The agreed timeline with timestamps and what was visible at each point
- ✓Contributing conditions, plural - not a single root cause
- ✓The near-misses and lucky catches that stopped it being worse
- ✓Three actions, each with a named owner, a date and a ticket reference
- ✓What was ruled out, so nobody re-litigates it in three weeks
- ✓No individual names attached to mistakes, anywhere in the document
- ○Published where a neighbouring team could stumble on it and learn
How to tell whether your reviews are working
The obvious measure - fewer incidents - is too noisy and too slow to steer by. Incident counts move with traffic, launches, hiring and luck, and a quarter of quiet can mean your reviews are working or simply that you shipped less.
Use three better signals instead. First, near-miss reporting: are people raising the things that nearly broke, unprompted, between incidents? That is the clearest evidence that the review process feels safe, because reporting a near-miss has no upside for the individual and only happens where the culture pays for it. Second, action completion within a month. Third, whether the same contributing condition appears in two reviews in a quarter - a repeat means the previous action addressed the symptom.
Add a twenty-second anonymous close to every review and track it over time. One question, five-point scale: did anything go unsaid today? A run of honest sixes and sevens out of ten is far more informative than a room that nods and leaves, and the trend tells you whether the format is earning trust or quietly losing it.
None of this requires a maturity model or a new process document. Run the same hour, in the same order, with the same question set, for six incidents in a row. By the fourth one the team will be rewriting each other's blaming questions before you have to, and that is the point at which the reviews start paying for themselves.
Frequently asked questions
What is a blameless post-incident review?
It is a structured session held a few days after an incident, in which the group reconstructs what happened and works out which conditions made the failure likely, without attributing it to an individual. The working assumption is that everyone acted sensibly with the information visible to them, so anything that looks careless in hindsight is treated as evidence about the system rather than about the person. The output is a timeline, a set of contributing conditions and a small number of owned actions.
How long should a post-incident review take?
Sixty minutes for a single-service incident with a clear beginning and end, ninety for a multi-team outage that ran for several hours, and a second shorter session a week later for anything that affected customers, money or data. The length matters less than the internal ratio: roughly a quarter of the time reading the timeline, a quarter in silent written contribution, a quarter in discussion and the remainder on actions. Most teams spend the whole hour on open discussion, which is why one voice ends up shaping the conclusion.
Who should attend an incident review?
The people who were directly involved, the people who own the affected systems, and one or two who will genuinely learn something. Eight to twelve is a good ceiling. Above about fifteen the session turns into a briefing: contributions concentrate in three voices and the engineer who spotted something odd at minute two says nothing. Invite adjacent teams when the incident crossed a boundary, and brief them on the blameless format beforehand so they do not arrive expecting a tribunal.
Should post-incident reviews be anonymous?
Make the written contribution phase anonymous when the incident touched customers, revenue or data, or when anyone in the room could reasonably worry about how it reflects on them. Keep the actions named, because owners have to be visible. Anonymity in the generation phase costs you nothing and buys you the contributions people would otherwise self-censor - the alert everyone ignores, the runbook nobody trusts, the process step that is routinely skipped because it does not work.
How do you stop a review turning into blame?
Three moves. State the boundary at the start - performance conversations happen elsewhere, and nothing here is used as evidence. Rewrite questions live, turning anything aimed at a person into a question about what was visible at that moment. And redirect the first blaming comment immediately with a single calm sentence, because the room reads your silence as permission. Self-blame needs the same treatment: accept it once, then ask what would have made the missed check automatic.
What should the incident write-up include?
Customer impact in plain language, the agreed timeline with what was visible at each point, contributing conditions in the plural rather than a single root cause, the near-misses that stopped it being worse, what was explicitly ruled out, and three actions with named owners and dates. Use roles instead of names throughout. Publish it somewhere a team two doors away can find it, since the same broken defaults are usually sitting in several other services.
Keep reading
- running a retrospective that changes something — The recurring cousin of the incident review, with the same problem of unowned actions.
- the evidence behind psychological safety — Why near-miss reporting is the signal that tells you the format is trusted.
- collecting anonymous feedback without losing trust — For the written contribution phase when the incident touched customers or data.
- facilitating difficult conversations in a group — Handling the moment a review turns tense despite the ground rules.
- run your incident review with Session Flo — Anonymous collection, ranking and a shareable recap from one room code.