Developer RunbooksObsidian & Markdown100% Free & Open Source
Production Incident Response & Blameless Post-Mortem Runbook
Engineering runbook for production outages: severity matrix (SEV-0 to SEV-3), incident commander protocols, status page communication templates, and blameless retrospective framework.
Designed For
DevOps Engineers, Site Reliability Engineers (SRE), CTOs, and Tech Leads
Companion OpenTools Utility
Developer Advanced WorkbenchOpen Tool
What's Included in this Template
SEV-0 through SEV-3 Severity Classification Matrix with response SLAs
Incident Commander & Communications Lead role definitions
Pre-written Status Page public communication snippets (Investigating, Identified, Monitoring, Resolved)
5-Whys Root Cause Analysis (RCA) and preventive action item tracker
Access & Download Template
Live Preview & Code InspectorZero Server Uploads
# Production Incident Response & Post-Mortem Runbook
Standard operating procedures for managing, mitigating, and documenting production service disruptions.
## 🚨 Incident Severity Classification Matrix
| Severity | Definition | Target Response Time | Update Cadence |
| :--- | :--- | :--- | :--- |
| **SEV-0 (Catastrophic)** | Total site outage, data loss risk, or active security breach | **< 5 minutes** | Every 15 minutes |
| **SEV-1 (Critical)** | Core workflow broken for > 20% of users (e.g. checkout / auth down) | **< 15 minutes** | Every 30 minutes |
| **SEV-2 (Major)** | Non-critical feature degraded; acceptable workaround exists | **< 1 hour** | Every 2 hours |
| **SEV-3 (Minor)** | Cosmetic bug, minor latency spike, single-user issue | **< 24 hours** | On resolution |
---
## 📢 Public Status Page Templates
### Update 1: Investigating
> "We are currently investigating reports of degraded performance affecting [Service Name]. Our engineering team is actively diagnosing the root cause. Next update in 15 minutes."
### Update 2: Identified
> "We have identified the root cause related to [Database Query Latency / Upstream API Provider] and are applying a mitigation patch now."
### Update 3: Monitoring & Resolved
> "The fix has been deployed and all system metrics have returned to nominal operating levels. We are continuing to monitor telemetry closely."
Frequently Asked Questions (FAQ)
Why is a blameless culture important for post-mortems?
Blameless retrospectives focus on systemic failures and process improvements rather than assigning individual fault, encouraging open and honest engineering disclosures.
More in Developer Runbooks
Obsidian & Markdown
Engineering Standard Operating Procedures (SOPs) Starter Kit
OperationsView