Release & Deployment
Estimated time: 15–30 min

Incident Response Checklist

A high-pressure operational runbook for on-call engineers to triage, communicate, mitigate, and resolve production outages.

#Incident Response#Outage#On-Call#DevOps#SRE#Mitigation#Postmortem#Production Failure

Live Incident Timeline & War Room Scratchpad

Phase Preset:
Progress:0%0 of 24 completed

Phase 1: Triage & Initial Assessment (SEV Classification)

Acknowledge alert, classify severity level, and establish leadership bridge.

0/4 checks
Acknowledge incident alert in PagerDuty / Opsgenie (<5 min)BlockingOn-Call Engineer
Classify Incident Severity Level (SEV1 / SEV2 / SEV3)BlockingIncident Commander
Designate Incident Commander (IC) & Communications LeadBlockingOn-Call Engineer
Open dedicated Incident War Room (Slack channel / Zoom bridge)BlockingIncident Commander

Phase 2: Communication & Stakeholder Alignment

Keep internal teams and affected customers updated with clear status notes.

0/3 checks
Update Public Status Page to 'Investigating'BlockingComms Lead
Notify internal leadership & customer support teamsImportantComms Lead
Establish periodic status update cadences (Every 15-30 min)ImportantComms Lead

Phase 3: Mitigation & Blast Radius Containment

STOP THE BLEEDING FIRST. Focus on restoring service before diagnosing root cause.

0/5 checks
Execute Deployment Rollback if recent code release occurredBlockingIncident Commander
Enable Feature Flag Killswitches for non-critical servicesBlockingIncident Commander
Apply rate limiting or API traffic sheddingBlockingDevOps
Trigger database failover to read-replica if primary corruptedBlockingDBA
Irreversible actions evaluated for blast radius impactBlockingIncident Commander

Phase 4: Investigation & Root Cause Diagnosis

Analyze metrics, trace logs, and formulate hypotheses while service is stabilized.

0/4 checks
Inspect APM golden signals (Latency, Traffic, Errors, Saturation)ImportantOn-Call Engineer
Query log streams for unhandled exceptions & OOM killsImportantOn-Call Engineer
Audit recent Git commits, DDL migrations & config changesImportantOn-Call Engineer
Test diagnostic hypotheses in isolated environmentImportantOn-Call Engineer

Phase 5: Resolution & Live Verification

Confirm baseline health metrics, resolve status page, and close war room.

0/4 checks
Confirm error rates & latency return to normal baselinesBlockingIncident Commander
Execute end-to-end production smoke testBlockingQA / Dev
Update Public Status Page to 'Resolved'BlockingComms Lead
Close Incident War Room bridgeImportantIncident Commander

Phase 6: Postmortem & Preventive Action Items

Conduct blameless postmortem, capture timeline, and file engineering debt fixes.

0/4 checks
Preserve timestamped incident timeline & log transcriptsBlockingIncident Commander
Schedule Blameless Postmortem meeting within 48 hoursBlockingEngineering Manager
File engineering remediation tickets for root cause fixBlockingIncident Commander
Update automated monitoring thresholds to catch recurrence earlierImportantDevOps
Golden Rule of Incident Response

Mitigation Comes Before Root Cause Analysis

During an active SEV1 or SEV2 outage, your sole job is to restore normal operations as fast as possible. Do not spend 45 minutes searching for the underlying root cause while users experience 100% downtime.

Phase A: Mitigate First (Stop the Bleeding)

Roll back recent deployments, enable feature flag killswitches, shed non-critical background traffic, or scale up server instances to bring error rates back to zero.

Phase B: Diagnose Root Cause Later

Once production is stable and baseline health is confirmed, inspect log streams, APM traces, and code diffs in a calm environment to understand why the failure occurred.

Operational Guardrails

High-Risk Irreversible Actions Warning Matrix

Operating under high pressure often induces panic. Never execute irreversible actions without explicit secondary sign-off from the Incident Commander:

1. Deleting Data or Truncating Database Tables

Never execute un-backed-up DELETE or TRUNCATE statements on production databases to clear corrupted records. Always export a snapshot first.

2. Forcing Database Failover Without Sync Verification

Forcing primary database failover when replication lag is high can cause permanent data loss for recent transactions. Check replica lag before failover.

3. Purging Persistent Message Queues

Purging RabbitMQ or SQS dead-letter queues discards customer webhooks and orders permanently. Re-route or pause consumer workers instead.

4. Disabling Security Controls or Auth Middleware

Bypassing authentication or rate-limiting middleware to resolve high latency exposes your production environment to immediate cyber attacks.

Postmortem Culture

The Blameless Postmortem Culture

Incidents are systemic failures, not individual engineering mistakes. Hold a blameless postmortem meeting within 48 hours focusing on process guardrails:

1. Timestamped Timeline

Record exact UTC timestamps for when the incident started, when alerts fired, when the IC responded, when mitigation was applied, and when normal health returned.

2. 5-Whys Analysis

Ask "Why?" five times to drill down to root systemic causes (e.g. "Why did the query fail? Missing index. Why was missing index not caught? Staging lacked production row volume.").

3. Action Item Owners

File Jira/GitHub tickets for engineering debt, automated test shims, and monitoring threshold updates. Assign explicit owners and target completion dates.

When to Use This Checklist

  • During active production outages, high 5xx error rate spikes, or database downtime.
  • When responding to PagerDuty / Opsgenie automated SEV1 or SEV2 alerts.
  • To guide on-call engineers through triage, mitigation, communication, and resolution under pressure.
  • During post-incident reviews to record timeline events and schedule blameless postmortems.

Common Pitfalls to Avoid

  • Wasting 45 minutes searching for root cause while users experience 100% outage, instead of mitigating immediately via rollback or traffic shedding.
  • Executing irreversible actions (deleting data, truncating DB tables, clearing queues) without evaluating blast radius.
  • Failing to update internal stakeholders and customer status pages, causing flood of duplicate support tickets.
  • Not recording timestamped timeline events during the incident, making postmortem analysis inaccurate.
  • Blaming individual engineers in postmortems instead of fixing systemic guardrails and automated testing shims.