Incident Response Checklist
A high-pressure operational runbook for on-call engineers to triage, communicate, mitigate, and resolve production outages.
Live Incident Timeline & War Room Scratchpad
Phase 1: Triage & Initial Assessment (SEV Classification)
Acknowledge alert, classify severity level, and establish leadership bridge.
Phase 2: Communication & Stakeholder Alignment
Keep internal teams and affected customers updated with clear status notes.
Phase 3: Mitigation & Blast Radius Containment
STOP THE BLEEDING FIRST. Focus on restoring service before diagnosing root cause.
Phase 4: Investigation & Root Cause Diagnosis
Analyze metrics, trace logs, and formulate hypotheses while service is stabilized.
Phase 5: Resolution & Live Verification
Confirm baseline health metrics, resolve status page, and close war room.
Phase 6: Postmortem & Preventive Action Items
Conduct blameless postmortem, capture timeline, and file engineering debt fixes.
Mitigation Comes Before Root Cause Analysis
During an active SEV1 or SEV2 outage, your sole job is to restore normal operations as fast as possible. Do not spend 45 minutes searching for the underlying root cause while users experience 100% downtime.
Phase A: Mitigate First (Stop the Bleeding)
Roll back recent deployments, enable feature flag killswitches, shed non-critical background traffic, or scale up server instances to bring error rates back to zero.
Phase B: Diagnose Root Cause Later
Once production is stable and baseline health is confirmed, inspect log streams, APM traces, and code diffs in a calm environment to understand why the failure occurred.
High-Risk Irreversible Actions Warning Matrix
Operating under high pressure often induces panic. Never execute irreversible actions without explicit secondary sign-off from the Incident Commander:
1. Deleting Data or Truncating Database Tables
Never execute un-backed-up DELETE or TRUNCATE statements on production databases to clear corrupted records. Always export a snapshot first.
2. Forcing Database Failover Without Sync Verification
Forcing primary database failover when replication lag is high can cause permanent data loss for recent transactions. Check replica lag before failover.
3. Purging Persistent Message Queues
Purging RabbitMQ or SQS dead-letter queues discards customer webhooks and orders permanently. Re-route or pause consumer workers instead.
4. Disabling Security Controls or Auth Middleware
Bypassing authentication or rate-limiting middleware to resolve high latency exposes your production environment to immediate cyber attacks.
The Blameless Postmortem Culture
Incidents are systemic failures, not individual engineering mistakes. Hold a blameless postmortem meeting within 48 hours focusing on process guardrails:
1. Timestamped Timeline
Record exact UTC timestamps for when the incident started, when alerts fired, when the IC responded, when mitigation was applied, and when normal health returned.
2. 5-Whys Analysis
Ask "Why?" five times to drill down to root systemic causes (e.g. "Why did the query fail? Missing index. Why was missing index not caught? Staging lacked production row volume.").
3. Action Item Owners
File Jira/GitHub tickets for engineering debt, automated test shims, and monitoring threshold updates. Assign explicit owners and target completion dates.
When to Use This Checklist
- During active production outages, high 5xx error rate spikes, or database downtime.
- When responding to PagerDuty / Opsgenie automated SEV1 or SEV2 alerts.
- To guide on-call engineers through triage, mitigation, communication, and resolution under pressure.
- During post-incident reviews to record timeline events and schedule blameless postmortems.
Common Pitfalls to Avoid
- Wasting 45 minutes searching for root cause while users experience 100% outage, instead of mitigating immediately via rollback or traffic shedding.
- Executing irreversible actions (deleting data, truncating DB tables, clearing queues) without evaluating blast radius.
- Failing to update internal stakeholders and customer status pages, causing flood of duplicate support tickets.
- Not recording timestamped timeline events during the incident, making postmortem analysis inaccurate.
- Blaming individual engineers in postmortems instead of fixing systemic guardrails and automated testing shims.
Connected Workflows & Tools
Complementary prompts, agent skills, and interactive tools in SprintKit.
Production Deployment Checklist
Pre-flight deployment readiness and rollback gates for shipping code.
Database Migration Checklist
Safely execute, backfill, and recover from database schema changes.
Security Review Checklist
Audit auth, permissions, input validation, and security vulnerabilities.
PR Review Queue
Workflow tool for tracking and prioritizing pull request reviews.