Skip to content

Template Design

Postmortem Template Design

A structured postmortem template ensures consistency, clarity, and actionable insights across incidents. Below is a recommended template structure, including formatting examples and content guidance.


1. Summary

Purpose: Provide a high-level overview of the incident.
Content:
- Incident Name: [e.g., "Database Write Latency Spike - 2023-10-05"]
- Date/Time: [e.g., "2023-10-05 14:32 UTC"]
- Impact: [e.g., "95th percentile latency increased from 12ms to 450ms for 15 minutes"]
- SLI/SLO Affected: [e.g., "API latency SLO (99th percentile < 50ms) violated"]
- Severity: [e.g., "Severity 2 (Partial Outage)"]

Example:

Incident Name: Database Write Latency Spike - 2023-10-05
Date/Time: 2023-10-05 14:32 UTC
Impact: 95th percentile latency increased from 12ms to 450ms for 15 minutes
SLI/SLO Affected: API latency SLO (99th percentile < 50ms) violated
Severity: Severity 2 (Partial Outage)


2. Timeline

Purpose: Document the incident's progression chronologically.
Content:
- Use timestamps and key events (e.g., alerts triggered, manual interventions).
- Include system states (e.g., "Database replica lagged by 10 seconds").

Example Table:
| Time | Event | Owner |
|-------------------|--------------------------------------------|-----------------|
| 14:32 UTC | Alert: "Database latency > 200ms" triggered | Monitoring Team |
| 14:35 UTC | Manual intervention: Scale up read replicas | SRE Team |
| 14:40 UTC | Latency stabilizes at 150ms | Database Team |


3. Root Cause

Purpose: Identify the root cause using systemic analysis.
Content:
- Immediate Cause: [e.g., "Disk I/O bottleneck due to misconfigured storage class"]
- Underlying Cause: [e.g., "Lack of automated monitoring for disk utilization"]
- Root Cause: [e.g., "Insufficient testing of storage class changes in staging"]

Example:

Immediate Cause: Disk I/O bottleneck due to misconfigured storage class
Underlying Cause: Lack of automated monitoring for disk utilization
Root Cause: Insufficient testing of storage class changes in staging


4. Action Items

Purpose: Define corrective steps with ownership and deadlines.
Content:
- Use a table with columns: Action, Owner, Deadline, Status.
- Prioritize actions (e.g., "Critical", "High", "Medium").

Example Table:
| Action | Owner | Deadline | Status |
|----------------------------------------|-----------------|----------------|-------------|
| Implement disk utilization alerts | Monitoring Team | 2023-10-12 | In Progress |
| Add storage class testing to CI/CD | DevOps Team | 2023-10-15 | Not Started |


5. Lessons Learned

Purpose: Capture systemic improvements.
Content:
- Systemic Insight: [e.g., "Manual interventions should be automated for latency spikes"]
- Process Gaps: [e.g., "Staging environments must mirror production storage configurations"]
- Tooling Gaps: [e.g., "Need a centralized incident response playbook"]

Example:

Systemic Insight: Manual interventions should be automated for latency spikes
Process Gaps: Staging environments must mirror production storage configurations
Tooling Gaps: Need a centralized incident response playbook


6. Mitigation Steps

Purpose: Document temporary fixes applied during the incident.
Content:
- List steps taken to resolve the issue (e.g., "Scale up read replicas", "Switch to cached storage").

Example:

  • Scale up read replicas to distribute I/O load
  • Switch to cached storage class temporarily
  • Enable real-time disk utilization metrics

Key takeaways

  • Standardize structure: Use consistent sections (summary, timeline, root cause, etc.) to ensure clarity.
  • Focus on systems, not individuals: Frame root causes as process or tooling gaps.
  • Prioritize action items: Assign owners, deadlines, and track progress to prevent recurrence.
  • Leverage data: Use metrics, logs, and traces to validate analysis and inform fixes.