Postmortem: Incident title
Incident postmortem – timeline, mermaid cause chain
उपयोग कैसे करें: After any user-facing incident. Blameless: focus on systems, not people. Timeline in UTC.
प्रीव्यू
Postmortem: Incident title
| Field | Value |
|---|---|
| Incident ID | INC-YYYY-NNNN |
| Severity | SEV-1 / SEV-2 / SEV-3 |
| Duration | YYYY-MM-DD HH:mm - HH:mm (NN min) |
| Impact | What users experienced |
| Author | @you |
| Status | Draft / Final |
Summary
Two sentences: what broke, how many users affected, how it was resolved.
Timeline (all times UTC)
| Time | Event |
|---|---|
| 14:02 | Alert fires: elevated 5xx on /api |
| 14:08 | On-call acknowledges, begins investigation |
| 14:25 | Root cause identified: bad deploy v2.3.1 |
| 14:31 | Rollback initiated |
| 14:38 | Service recovered |
Root cause
Technical explanation of what actually failed. Include the contributing factors, not just the trigger.
flowchart TD
A[Deploy v2.3.1] --> B[Config flag missing]
B --> C[Null pointer on request path]
C --> D[5xx spike]
What went well
- Detection was fast (alert within 2 min)
- Rollback runbook worked as written
What went poorly
- Deploy pipeline did not catch the missing flag
- Runbook step 3 was outdated
Action items
| Action | Owner | Due | Priority |
|---|---|---|---|
| Add config validation to CI | @dev | YYYY-MM-DD | P0 |
| Update rollback runbook | @ops | YYYY-MM-DD | P1 |
| Add canary deploy stage | @dev | YYYY-MM-DD | P1 |
Lessons learned
The takeaway that changes how we build or operate going forward.