blameless-postmortem

Date: 2025-10-27 Incident Commander: Alice Smith (@alice) Severity: Sev2 Duration: 45 minutes (14:23 - 15:08 UTC) Impact: ~15,000 users experienced errors (10% of active users) Error Budget Impact: 15% of monthly budget consumed

Incident Postmortem: API High Error Rate

Date: 2025-10-27 Incident Commander: Alice Smith (@alice) Severity: Sev2 Duration: 45 minutes (14:23 - 15:08 UTC) Impact: ~15,000 users experienced errors (10% of active users) Error Budget Impact: 15% of monthly budget consumed

Executive Summary

The Payment API experienced elevated error rates (5-10%) for 45 minutes due to database connection pool exhaustion. A recent deployment introduced an N+1 query pattern that, combined with a 3x traffic spike from a marketing campaign, saturated database connections. Service was restored by rolling back the deployment. No data was lost or corrupted.

Root Cause: N+1 query pattern introduced in v2.3.4, triggered by unexpected traffic spike.

Key Learnings:

Timeline

All times in UTC.

14:20 - Deployment v2.3.4 completed successfully

14:23 - First alert fired: Error rate > 1%

14:25 - On-call engineer @alice acknowledged alert

14:27 - Traffic analysis revealed 3x normal load

14:30 - Identified database connection pool saturation

14:33 - Database pool exhausted

14:35 - Incident commander assigned (@alice)

14:37 - Root cause identified: N+1 query in v2.3.4

14:40 - Decision: Rollback to v2.3.3

14:42 - Rollback initiated

14:45 - Rollback complete (v2.3.3 deployed)

14:50 - Error rate returned to baseline

14:55 - Monitoring for stability

15:00 - Customer communication sent

15:08 - Incident declared resolved

Root Cause Analysis

Five Whys

Problem: API returned 500 errors

  1. Why? Database connection timeouts
  2. Why? Connection pool exhausted
  3. Why? Too many concurrent database queries
  4. Why? N+1 query pattern + high traffic
  5. Why? Code change not tested under load

Root Cause: Lack of load testing in CI/CD pipeline allowed inefficient query pattern to reach production.

Technical Details

Code Change (v2.3.4):

# BAD: N+1 query pattern
def get_products(category_id):
    products = db.query("SELECT * FROM products WHERE category_id = ?", category_id)
    for product in products:
        # N queries (one per product)
        product.reviews = db.query("SELECT * FROM reviews WHERE product_id = ?", product.id)
    return products

Should have been:

# GOOD: Single query with JOIN
def get_products(category_id):
    return db.query("""
        SELECT p.*, r.*
        FROM products p
        LEFT JOIN reviews r ON p.id = r.product_id
        WHERE p.category_id = ?
    """, category_id)

Impact Under Load:

Contributing Factors

  1. No load testing in CI/CD
  1. Marketing campaign not coordinated
  1. Insufficient query monitoring
  1. Manual deployment rollback
  1. Missing connection pool alerts

Impact Assessment

User Impact

System Impact

SLO Impact

Customer Sentiment

Resolution

Immediate Fix

Rolled back to v2.3.4, which restored service within 10 minutes of decision.

Permanent Fix

Fixed N+1 query in v2.3.5:

Action Items

Prevention (Stop it from happening)

| Priority | Action | Owner | Due Date | Ticket | |----------|--------|-------|----------|--------| | P0 | Add load testing to CI/CD pipeline | @bob | 2025-11-10 | ENG-1234 | | P0 | Implement query performance monitoring | @charlie | 2025-11-15 | ENG-1235 | | P1 | Add database connection pool alerts | @alice | 2025-11-08 | ENG-1236 | | P1 | Establish marketing-engineering coordination | @diana | 2025-11-20 | ENG-1237 | | P2 | Add query linter to catch N+1 patterns | @bob | 2025-11-25 | ENG-1238 |

Detection (Find it faster)

| Priority | Action | Owner | Due Date | Ticket | |----------|--------|-------|----------|--------| | P0 | Add query execution time to traces | @charlie | 2025-11-12 | ENG-1239 | | P1 | Reduce alert evaluation window (5m → 1m) | @alice | 2025-11-05 | ENG-1240 | | P1 | Add per-endpoint query count metrics | @bob | 2025-11-18 | ENG-1241 |

Response (Fix it faster)

| Priority | Action | Owner | Due Date | Ticket | |----------|--------|-------|----------|--------| | P0 | Implement auto-rollback on error threshold | @bob | 2025-11-18 | ENG-1242 | | P1 | Document connection pool troubleshooting | @alice | 2025-11-08 | ENG-1243 | | P2 | Create runbook for database saturation | @charlie | 2025-11-15 | ENG-1244 |

Process Improvements

| Priority | Action | Owner | Due Date | Ticket | |----------|--------|-------|----------|--------| | P1 | Require load testing sign-off for deploys | @manager | 2025-11-12 | ENG-1245 | | P1 | Create marketing-engineering notification process | @product | 2025-11-10 | ENG-1246 | | P2 | Add capacity planning review (quarterly) | @sre | 2025-12-01 | ENG-1247 |

Lessons Learned

What Went Well ✓

  1. Quick detection (3 minutes)
  1. Clear incident command
  1. Decisive rollback decision
  1. No data loss
  1. Good customer communication

What Went Poorly ✗

  1. No load testing caught the issue
  1. Manual rollback took 10 minutes
  1. Marketing campaign not communicated
  1. Missing database performance monitoring
  1. No connection pool alerts

Where We Got Lucky 🍀

  1. Incident during business hours
  1. Simple rollback path available
  1. Database connections recovered quickly
  1. No cascading failures
  1. User perception

Metrics and Data

Error Budget

Performance Impact

| Metric | Baseline | During Incident | Peak | |--------|----------|-----------------|------| | Error rate | < 0.1% | 5-10% | 10% | | Latency (p95) | 50ms | 200-500ms | 500ms | | Latency (p99) | 200ms | 1000ms+ | 2000ms | | DB connections | 30/50 | 50/50 | 50/50 (exhausted) | | Query time | 50ms | 300ms | 600ms |

Business Impact

Supporting Information

References

Attendees

Follow-up


Approval:

Document Version: 1.0 Status: Final