Brief description of what this alert means and why it exists.
Status: [Draft | Active | Archived] Last Updated: YYYY-MM-DD Owner: [Team/Person] Severity: [Critical | Warning | Info]
Brief description of what this alert means and why it exists.
What is being measured: Describe the metric or condition
Why it matters: Explain the business or technical impact
What the oncall engineer will observe:
Current Severity: [Critical | Warning | Info]
Escalate to Critical if:
Escalate to Manager if:
Check if alert is accurate:
# Query Prometheus to verify metric
curl 'http://prometheus:9090/api/v1/query?query=[METRIC_NAME]'
# Check Grafana dashboard
# URL: [Dashboard URL]
Expected: [What should you see?] Actual: [What are you seeing?]
# Check service health
kubectl get pods -n [namespace]
kubectl logs -f [pod-name] --tail=100
# Check dependencies
curl -I https://[dependency]/health
# Check recent deployments
kubectl rollout history deployment/[name]
# Check git commits
git log --since="1 hour ago" --oneline
# Check recent alerts
# Alertmanager URL: [URL]
Common causes:
Objective: Stop the bleeding, restore service
``bash # Command with explanation kubectl scale deployment/[name] --replicas=[N] `` Expected outcome: [What should happen] Verification: [How to verify it worked]
``bash # Command ``
Objective: Stable workaround while investigating root cause
Objective: Permanent resolution
If remediation makes things worse:
# Rollback deployment
kubectl rollout undo deployment/[name]
# Restore from backup
./restore_backup.sh [backup-id]
# Disable feature flag
curl -X POST https://[feature-flags]/disable/[flag]
| Stakeholder | When | How | |-------------|------|-----| | Team Lead | Immediately | Slack #[channel] | | Manager | If not resolved in 30min | Phone/Slack | | Customers | If impact confirmed | Status page | | Exec Team | If revenue impact | Email + Slack |
Initial notification (within 5 minutes):
[INCIDENT] [Service] - [Brief description]
Impact: [User-facing impact]
Started: [Time]
Status: Investigating
Next update: [Time]
Incident lead: [Name]
Slack channel: #incident-[ID]
Status update (every 30 minutes):
[UPDATE] [Service] - [Brief description]
Actions taken:
- [Action 1]
- [Action 2]
Current status: [Status]
Next update: [Time]
Resolution:
[RESOLVED] [Service] - [Brief description]
Resolution: [What fixed it]
Root cause: [Brief description]
Duration: [Total time]
Postmortem: [Link] (to be completed within 48h)
Verify service is healthy:
# Check 1: Service responding
curl https://[service]/health
# Check 2: Metrics recovered
# [Dashboard URL]
# Check 3: Error rate normal
# [Query URL]
# Check 4: Latency acceptable
# [Dashboard URL]
All checks passing? ✓ Incident resolved
[AlertName]: [Brief description][AlertName]: [Brief description]# View logs
kubectl logs -f [pod] --namespace=[ns] --tail=100
# Check resource usage
kubectl top pods -n [namespace]
# Describe pod
kubectl describe pod [pod] -n [namespace]
# Execute command in pod
kubectl exec -it [pod] -n [namespace] -- bash
# Port forward for debugging
kubectl port-forward [pod] 8080:8080 -n [namespace]
# Request rate
sum(rate(http_requests_total[5m])) by (service)
# Error rate
sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)
/
sum(rate(http_requests_total[5m])) by (service)
# Latency p95
histogram_quantile(0.95,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service)
)
| Date | Author | Change | |------|--------|--------| | YYYY-MM-DD | [Name] | Initial version | | YYYY-MM-DD | [Name] | Updated remediation steps |
Found an issue with this runbook? Incident didn't match the runbook?