Get 1 month of Premium free

Use the code at checkout

00Days
00Hours
00Mins
00Secs
Claim 1 month free

BlogOperations

How to Run a Post-Mortem After Something Goes Wrong

Build the timeline before you discuss causes. Establish what happened and when, using facts rather than recollections, then ask what made each decision reasonable at the time. Blame stops the flow of information you need, so a post-mortem that identifies a person has usually failed to find the cause.

Why blame breaks the process

The instinct after something goes wrong is to find out who did it. It feels like accountability and it destroys the thing you need most, which is an accurate account of what happened.

People who expect to be blamed do three things, all rational: they describe their own actions in the most favourable terms, they omit details that look bad, and they do not volunteer near misses. You end up with a tidy narrative that supports a conclusion someone had already reached, and the actual cause survives untouched to produce the next incident.

The working assumption

Everyone acted reasonably given what they knew at the time. If someone did something that looks obviously wrong in hindsight, the question is what made it look right in the moment. That is where the fixable cause lives.

This is not about being soft. A post-mortem that concludes "Priya deployed without checking" has produced no change: the next person will also deploy without checking, because nothing about the situation is different. One that concludes "there is no check between deploy and production, and the deploy button looks identical in both environments" has produced two fixable things.

When to run one

Not for everything. Three triggers are enough:

Something reached a customer. An outage, a wrong invoice, a missed deadline they noticed.

Something took far longer than expected to resolve. Duration is itself a finding, even where the outcome was fine.

A near miss that would have been serious. These are the cheapest reviews you will ever run, and the ones teams almost always skip.

Run it within a week. Sooner and people are still tired and defensive; later and everyone has converted their memory into a story.

The method

Ninety minutes. The order matters more than anything else here.

  1. Build the timeline before discussing anything

    Facts with timestamps, from messages, logs, emails and calendars rather than from memory. What happened, when, and who knew what at each point.

    Do this first and resist all analysis while doing it. Teams that discuss causes before establishing the timeline reliably build a timeline that supports the cause they already suspected.

    30 minutesFacts only, no interpretation
  2. Mark the moments where a different decision was available

    Points where someone chose, noticed, or could have noticed something. Not mistakes, just decision points.

    Usually three to five in any incident, and they are rarely where people expected before the timeline was built.

    10 minutes
  3. For each, ask what made that choice reasonable

    What information did they have? What did the interface show? What was the pressure? What had happened the last twenty times?

    This is the core of the method. The answers are the causes, and they are almost always about information, tooling, or process rather than about attention.

    30 minutesThe actual work
  4. Ask what made it take as long as it did to fix

    Detection, escalation, access, knowledge. A two-hour outage where ninety minutes was spent finding the person with the right login is a different problem from a technical one.

    Response time is frequently more improvable than prevention, and it is routinely ignored.

    10 minutesOften the bigger win
  5. Agree one or two changes, with an owner and a date

    Not eight. One completed change beats five identified ones, which is the same discipline that makes a [retrospective](/blog/how-to-run-a-retrospective) work.

    Pick changes that make the next incident less likely or less damaging, not ones that ask people to be more careful.

    10 minutesOne or two, owned

Finding the cause without hindsight bias

Two traps account for most bad post-mortems.

Hindsight bias. Once you know the outcome, the warning signs look obvious. They were not obvious at the time, mixed in with dozens of other signals that led nowhere. The corrective is to keep asking what else was happening simultaneously, and what the same signal had meant on previous occasions.

Stopping at human error. "Someone made a mistake" is where an investigation ends, not where it should. Push one level further every time: what allowed that mistake to have this consequence? Systems where a single human error causes serious damage are the finding.

Two questions that reliably find something useful:

"What did we think was true that was not?" Almost every incident contains a wrong belief about the state of something.

"Who would have needed to know, and how would they have found out?" This surfaces detection and communication gaps, which are usually easier to fix than the technical cause.

Making the actions happen

The failure mode is a well-run review producing a document nobody acts on.

One or two actions, with a named owner and a date. Recorded where the team already looks, not only in the post-mortem document. The mechanics are the same as meeting notes that turn into action items.

Check them at the next one. Start every post-mortem by reviewing whether the last one's actions happened. That single habit is what separates a team that learns from one that documents.

Write it up short and consistently. What happened, timeline, causes, actions. One page, same shape every time. The consistency is what lets you see patterns across six months, which is where the real findings are, and the format rules in internal docs people actually read apply.

Look across them quarterly. Individual incidents look unique. Six of them together usually show two or three recurring themes: a handover that keeps losing context, a system only one person understands, a step that is manual and shouldn't be. Those themes are worth more than any single review.

Where the pattern is that only one person could resolve something, that is the same single-route dependency covered in preparing the business for someone going on leave, and it is a risk whether or not anyone is on holiday. Where the pattern is customer-facing, our guide to handling customer escalations covers the response side that runs in parallel with the investigation.

Frequently asked questions

What is a blameless post-mortem?
A review that treats the outcome as a product of the system rather than of a person's failure. The assumption is that everyone acted reasonably given what they knew, and the useful question is what made a wrong action look right at the time.
How soon after an incident should you run a post-mortem?
Within a week, once the immediate problem is resolved. Sooner and people are still exhausted and defensive; much later and the details have faded into a story everyone has told themselves.
Who should attend?
The people who were involved, plus whoever can authorise the fixes. Keep it small. Observers change what people are willing to say, and the honesty is the entire value of the meeting.
What if someone genuinely made a mistake?
Then the question is what allowed a single mistake to cause this much damage. A system where one person's error produces a serious incident has a design problem, and that is the finding worth acting on.
Should post-mortems be written up and shared?
Yes, internally, in a consistent short format. The write-ups accumulate into the most honest record you have of how your business actually fails, and patterns emerge across them that no single review reveals.
Danish Khan

Danish Khan

CEO & Founder, Siela

Danish Khan is the CEO and founder of Siela, an AI-native workspace where teams and AI agents run CRM, meetings, tasks, and daily work together on one shared context layer.

Connect on LinkedIn

Published