After the incident: Running a post-mortem that actually changes security behavior
Services are back. The incident channel has gone quiet. People who spent days dealing with the breach are catching up on everything that stopped while they were responding.
Now, it's time for the post-mortem. Someone presents a timeline, the team discusses what went wrong, and the meeting ends with recommendations: improve awareness, strengthen controls, and update procedures.
A month later, those recommendations may still be recommendations.
A useful post-mortem closes that gap. It explains why the incident was possible, identifies what needs to change and establishes how the organization will know the changes worked. Its value becomes visible in everyday decisions after the urgency has faded.
Bring the right people, with something to review
Arrange the review once the immediate pressure has eased and participants can contribute meaningfully. Capture observations while memories are fresh; unresolved investigation findings can be marked as provisional and updated later.
Invite people who understand the affected technology and the work around it. This may include security, IT operations, the service owner, and someone from the affected business team. A security team can identify an unsafe workaround without knowing why employees depend on it.
Circulate a short account beforehand: what happened, the business impact, the main decisions and the questions still unanswered. Use the meeting to examine those questions rather than read the incident record aloud.
Choose a facilitator who can challenge assumptions without turning the discussion into a defense of individual decisions. Give quieter participants a way to contribute in writing.
Ask why the decision made sense at the time
The employee clicked the link describes an action. It does not explain why the organization was vulnerable to its consequences.
What made the request convincing? Was the employee following a familiar process? Could they verify it easily? What protection should have limited the damage after the click?
Apply the same questions to technical decisions. If a patch was deferred, examine the reason: a failed compatibility test, an unclear service owner, or an exception that never expired.
Google’s guidance on blameless post-mortems recommends examining contributing conditions and the information available to people when they acted. That approach helps surface problems that accusations can obscure. [1]
Accountability still matters. Record decisions accurately and assign responsibility for improvements. However, be more careful is a weak corrective action when the same confusing process remains in place.
Look beyond the first thing that failed
Avoid forcing the incident into one convenient root cause. Several conditions may have combined to make it possible or increase its impact.
Separate the discussion into three questions:
What allowed the incident to begin?
What allowed its consequences to grow?
What made the organization’s handling of it harder?
An exposed vulnerability might answer the first question. Excessive access might answer the second. An inaccurate dependency record might answer the third. Fixing the vulnerability alone leaves the other weaknesses available for a different attack.
Also record what worked. If an employee questioned an unusual request or a routine review uncovered unexpected access, identify what supported that behavior. Successful practices deserve deliberate reinforcement.
For every explanation, ask what evidence supports it. Where evidence is missing, record the uncertainty instead of turning a plausible theory into a fact.
Replace recommendations with work someone can finish
Improve access management is too broad to assign, schedule or verify.
A stronger action names the affected scope and the expected result: review privileged membership for the affected application, remove access without a current business justification, and have the application owner approve the remaining membership.
Each action should have a responsible party, due date, priority, and evidence required for closure. Google’s post-mortem workbook similarly emphasizes concrete, tracked follow-up actions. [2]
Use a simple action record:
Finding | Assigned change | Evidence of completion |
Temporary access remained active | Introduce expiry and owner review for temporary grants | Sample grants expire as intended |
Staff could not find the approved process | Put instructions where the task is performed | Intended users can locate and follow them |
A security exception had no review date | Assign an owner and expiry to each active exception | Exception register reviewed and approved |
Agree on a manageable set of priorities. If a change needs funding or another team’s work, record that dependency. An unfunded recommendation should not appear as an accepted delivery commitment.
Make the safer behavior easier
Training can explain what people should do. The workflow determines how practical that choice is.
If employees bypassed an approved tool because access took weeks, another reminder to use it leaves the obstacle untouched. Give someone responsibility for improving that access process.
If a payment-verification step was skipped, examine whether contact details were available and whether the procedure worked during busy periods. Make the expected action specific enough to follow under ordinary working pressure.
Share lessons in a form each audience can use. Employees may need a brief explanation of a changed process. Administrators may need an updated configuration standard. Managers may need to approve time or resources.
Avoid distributing sensitive incident details more widely than necessary. People need enough context to understand the change and apply it.
Close actions with evidence, then check again
Implementation and effectiveness are separate questions.
A policy can be published without changing behavior. A technical control can be enabled without covering every intended system. Before closing an action, check its result against the condition identified in the review.
For a process change, ask representative users to complete the task. For a technical change, use an authorized test appropriate to the control and environment. Record exceptions and assign follow-up work.
Set a follow-up date when the action is agreed. Review whether the improvement is still working, whether people have developed new workarounds and whether the same finding has appeared in another incident.
The absence of another breach is not enough to prove success. Look for direct evidence that the original weakness has been reduced.
A post-mortem is useful when the next person facing the same decision has better information, a clearer process or a stronger control. The meeting produces the findings. Following through changes security behavior.
Sources
Google SRE: Postmortem Culture—Learning from Failure — blameless analysis and contributing conditions.
Google SRE Workbook: Postmortem Culture—Learning from Failure — actionable follow-up and post-mortem practices.