markdowneditor

Postmortem: Incident title

Incident postmortem – timeline, mermaid cause chain

Software & EngineeringStandardpostmortemdiagram

How to use: After any user-facing incident. Blameless: focus on systems, not people. Timeline in UTC.

Preview

Postmortem: Incident title

FieldValue
Incident IDINC-YYYY-NNNN
SeveritySEV-1 / SEV-2 / SEV-3
DurationYYYY-MM-DD HH:mm - HH:mm (NN min)
ImpactWhat users experienced
Author@you
StatusDraft / Final

Summary

Two sentences: what broke, how many users affected, how it was resolved.

Timeline (all times UTC)

TimeEvent
14:02Alert fires: elevated 5xx on /api
14:08On-call acknowledges, begins investigation
14:25Root cause identified: bad deploy v2.3.1
14:31Rollback initiated
14:38Service recovered

Root cause

Technical explanation of what actually failed. Include the contributing factors, not just the trigger.

flowchart TD
    A[Deploy v2.3.1] --> B[Config flag missing]
    B --> C[Null pointer on request path]
    C --> D[5xx spike]

What went well

  • Detection was fast (alert within 2 min)
  • Rollback runbook worked as written

What went poorly

  • Deploy pipeline did not catch the missing flag
  • Runbook step 3 was outdated

Action items

ActionOwnerDuePriority
Add config validation to CI@devYYYY-MM-DDP0
Update rollback runbook@opsYYYY-MM-DDP1
Add canary deploy stage@devYYYY-MM-DDP1

Lessons learned

The takeaway that changes how we build or operate going forward.

Related templates