This document provides a complete walkthrough of responding to a SEV-1 (critical) incident from detection to postmortem.
This document provides a complete walkthrough of responding to a SEV-1 (critical) incident from detection to postmortem.
Incident: High error rate (35%) in payment processing API Detection: Automated monitoring alert at 14:25 UTC Impact: 50% of payment transactions failing, affecting thousands of users
PagerDuty Alert:
CRITICAL: High Error Rate in Payment API
Error rate: 35% (threshold: 5%)
Dashboard: https://grafana.example.com/d/payments
Runbook: https://runbooks.example.com/payments/high-error-rate
On-call engineer (Alice) receives page
Alice acknowledges alert and checks dashboard:
# Quick checks
curl https://api.example.com/health
# Returns: {"status": "degraded", "payment_api": "unhealthy"}
# Check recent deployments
kubectl rollout history deployment/payment-service | tail -5
# Shows: deployment v2.3.1 at 14:15 UTC (10 minutes ago)
# Check error logs
kubectl logs -l app=payment-service --since=15m | grep ERROR | head -20
# Shows: database connection errors
Assessment:
Alice determines this is SEV-1:
Actions:
# 1. Create incident in tracking system
./create_incident.py create \
--title "High error rate in payment API" \
--severity SEV-1 \
--impact "50% of payment transactions failing" \
--service payment-api \
--detected-by monitoring
# Output: Created incident: INC-20251027-ABC123
# 2. Create PagerDuty incident
python pagerduty-integration.py create \
--incident-id INC-20251027-ABC123 \
--service-id PXXXXXX \
--title "SEV-1: Payment API High Error Rate"
# 3. Create Slack war room
python slack-incident-bot.py create-war-room \
--incident-id INC-20251027-ABC123 \
--title "High error rate in payment API" \
--severity SEV-1 \
--ic-user-id U123ABC # Alice
# Output: Created war room: #incident-20251027-abc123
Alice posts in #incident-20251027-abc123:
@here SEV-1 Incident Declared
INC-20251027-ABC123: High error rate in payment API
IMPACT:
- 35% error rate in payment transactions
- Affecting 50% of users attempting payments
- Started: ~14:20 UTC (correlates with deployment v2.3.1)
CURRENT STATUS: Investigating
TEAM:
- IC: @alice
- On-call: @alice (same person, need backup)
FIRST ACTIONS:
- Assessing rollback vs. fix
- Need database team to check connection pool
- Communications: notify support team NOW
Next update in 15 minutes.
# Page database specialist
pagerduty incident create \
--service database-service \
--title "DB connection errors - Payment incident" \
--urgency high
# Page engineering manager for IC support
pagerduty incident create \
--escalation-policy eng-management \
--title "SEV-1 needs IC support"
# Invite to war room
python slack-incident-bot.py invite \
--channel incident-20251027-abc123 \
--users U234BCD,U345CDE # Bob (DB), Charlie (Mgr)
Communications Lead (auto-assigned or manually):
Subject: [SEV-1] Payment Processing Issue
We are currently experiencing elevated error rates in payment processing.
IMPACT:
- Some payment transactions may fail
- Users may see error messages during checkout
- Started at: 14:20 UTC
STATUS:
- Issue identified and team actively investigating
- Working on mitigation
NEXT UPDATE:
- Will provide update in 30 minutes
- Status page: https://status.example.com
We apologize for the inconvenience.
War room participants:
Charlie (Mgr) posts:
[14:36] @charlie Taking formal IC role. @alice moving to tech lead.
Roles:
- IC: @charlie
- Tech Lead: @alice
- Database SME: @bob
- Communications: @diana
- Scribe: @eve
@alice, what's your recommendation? Rollback or investigate?
Timeline in war room:
[14:37] @alice [OBSERVATION] Deployment v2.3.1 correlates exactly with error spike
[14:38] @alice [OBSERVATION] Logs show "connection pool exhausted" errors
[14:38] @bob [OBSERVATION] DB connection count at 95/100 (near max)
[14:39] @bob [OBSERVATION] Queries look normal, but higher volume than usual
[14:40] @charlie [DECISION] Rollback deployment v2.3.1 to v2.3.0
[14:40] @alice [ACTION] Starting rollback now
Rollback execution:
# 1. Rollback deployment
kubectl rollout undo deployment/payment-service
# 2. Verify rollback
kubectl rollout status deployment/payment-service
# Output: successfully rolled out
# 3. Monitor metrics
watch -n 5 'curl -s https://api.example.com/metrics | grep payment_error_rate'
[14:42] @alice [ACTION] Rollback complete, v2.3.0 now running
[14:43] @alice [OBSERVATION] Error rate still at 30%
[14:43] @charlie [DECISION] Rollback didn't help immediately, monitor for 5 minutes
[14:45] @alice [OBSERVATION] Error rate declining: 22%
[14:46] @bob [OBSERVATION] DB connection count dropping: 75/100
[14:47] @alice [OBSERVATION] Error rate at 10%
[14:48] @alice [OBSERVATION] Error rate at 3%
[14:48] @bob [RECOMMENDATION] Increase connection pool to prevent recurrence
[14:49] @charlie [DECISION] Proceed with pool increase
[14:49] @bob [ACTION] Increasing pool from 20 to 50
# Increase connection pool
kubectl set env deployment/payment-service DB_POOL_SIZE=50
# Verify
kubectl get deployment payment-service -o jsonpath='{.spec.template.spec.containers[0].env[?(@.name=="DB_POOL_SIZE")].value}'
# Output: 50
[14:50] @alice [OBSERVATION] Error rate back to baseline: 0.5%
[14:50] @charlie [DECISION] Monitoring for 15 minutes before declaring resolved
Subject: [SEV-1] Update #1 - Payment Processing Issue
UPDATE #1 - 14:50 UTC
CURRENT STATUS: Mitigated
ACTIONS TAKEN:
- 14:40 UTC: Rolled back recent deployment
- 14:49 UTC: Increased database connection pool capacity
- Error rate decreased from 35% to 0.5% (baseline)
CURRENT EFFORTS:
- Monitoring for stability before declaring resolved
- Investigating root cause
NEXT UPDATE: 15:20 UTC (or when resolved)
[14:52] @alice [OBSERVATION] All metrics stable for 2 minutes
[14:55] @alice [OBSERVATION] No new errors in logs
[14:58] @bob [OBSERVATION] DB connection pool at healthy 15/50 utilization
[15:00] @alice [OBSERVATION] Payment success rate back to 99.5% (normal)
[15:02] @alice [OBSERVATION] No customer complaints in last 10 minutes
[15:05] @alice [OBSERVATION] 15 minutes stable, all criteria met
Charlie (IC) verifies resolution criteria:
Resolution Criteria:
✓ Error rate back to baseline (<0.5%)
✓ Latency within SLO (p95 < 500ms)
✓ Throughput recovered to normal
✓ All health checks passing
✓ No active alerts
✓ User-facing functionality verified
✓ Support ticket volume normal
✓ No customer reports for 15+ minutes
✓ System stable for 15+ minutes
✓ Related systems healthy
✓ Dependencies verified operational
ALL CRITERIA MET
[15:07] @charlie [DECISION] Declaring incident RESOLVED
[15:07] @charlie [NOTIFICATION] Incident INC-20251027-ABC123 is RESOLVED
[15:07] @charlie [ACTION] @diana send resolution notification
[15:08] @charlie [ACTION] @eve schedule postmortem for tomorrow 10:00 UTC
Update systems:
# Update incident status
./create_incident.py update INC-20251027-ABC123 \
--status resolved \
--event "Incident resolved" \
--author charlie
# Resolve PagerDuty incident
python pagerduty-integration.py resolve \
--incident-id PXXXXXX \
--resolution "Rolled back deployment and increased connection pool"
# Update Slack
python slack-incident-bot.py post-resolution \
--channel incident-20251027-abc123 \
--incident-id INC-20251027-ABC123 \
--duration "45 minutes" \
--resolution "Rolled back deployment v2.3.1 and increased DB connection pool to 50"
Subject: [RESOLVED] Payment Processing Issue Resolved
The incident affecting payment processing has been RESOLVED.
SUMMARY:
- Started: 14:20 UTC
- Resolved: 15:07 UTC
- Duration: 47 minutes
- Impact: Elevated error rates in payment transactions
RESOLUTION:
- Rolled back deployment v2.3.1
- Increased database connection pool capacity
- Service is now operating normally
- Monitoring for continued stability
ROOT CAUSE:
- Preliminary: Deployment introduced query changes that exhausted DB connection pool
- Full postmortem will be published within 48 hours
NO ACTION REQUIRED from users. All pending payments will be automatically retried.
We apologize for any inconvenience this may have caused.
[15:12] @charlie Thank you team for excellent response:
- @alice: Quick diagnosis and rollback
- @bob: Fast DB analysis and mitigation
- @diana: Clear customer communications
- @eve: Great timeline documentation
MTTR: 47 minutes from detection to resolution
Postmortem: Tomorrow 10:00 UTC in this channel
This war room will stay open for documentation.
Actions:
# Compile timeline
python slack-incident-bot.py export-timeline \
--channel incident-20251027-abc123 \
--output timeline.json
# Gather metrics
curl "https://grafana.example.com/api/dashboards/uid/payments" > metrics.json
# Take screenshots of key graphs
# (manual or automated)
# Begin postmortem draft
./generate_postmortem.py \
--incident-id INC-20251027-ABC123 \
--authors alice,bob,charlie \
--template deployment \
--output postmortems/INC-20251027-ABC123.md
Draft sections completed:
Key findings:
Agenda (1 hour):
10:00-10:05: Read executive summary
10:05-10:15: Walk through timeline
10:15-10:25: What went well
10:25-10:40: What went poorly (blameless!)
10:40-10:55: Action items
10:55-11:00: Assign owners and deadlines
Action Items Defined:
Prevent Recurrence:
Improve Detection:
Improve Response:
Improve Processes:
Published to:
Follow-up scheduled:
This workflow can be adapted for other SEV-1 incidents by: