This alert fires when memory usage exceeds 90% for more than 5 minutes on any node.
Status: Active Last Updated: 2025-10-27 Owner: Platform Team Severity: Critical
This alert fires when memory usage exceeds 90% for more than 5 minutes on any node.
What is being measured: node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes
Why it matters: High memory usage can lead to:
What you'll observe:
Current Severity: Critical (if > 95%), Warning (if > 90%)
Escalate to Critical if:
Escalate to Manager if:
# SSH to affected node
ssh [node-ip]
# Check current memory usage
free -h
# Example output:
# total used free shared buff/cache available
# Mem: 15Gi 14Gi 100Mi 50Mi 900Mi 800Mi
# Swap: 2.0Gi 1.5Gi 500Mi
# Check memory usage percentage
echo "scale=2; $(grep MemAvailable /proc/meminfo | awk '{print $2}') / $(grep MemTotal /proc/meminfo | awk '{print $2}') * 100" | bc
# View top memory consumers
top -o %MEM
# Or use htop (better visualization)
htop
Expected: Memory available > 10% (> 1.5GB on 16GB node) Actual: Memory available < 10%
# Top 10 processes by memory (Linux)
ps aux --sort=-%mem | head -11
# Detailed per-process memory breakdown
ps aux | awk '{print $6/1024 " MB\t" $11}' | sort -n
# Container memory usage (Docker)
docker stats --no-stream --format "table {{.Container}}\t{{.MemUsage}}"
# Kubernetes pod memory usage
kubectl top pods --all-namespaces --sort-by=memory
# Check for memory leaks in specific process
pmap -x [PID] | tail -1
# Monitor process memory over time
watch -n 5 'ps aux | grep [process-name]'
# Check for zombie processes
ps aux | grep defunct
# View kernel memory slab usage
sudo slabtop
# Check page cache usage
cat /proc/meminfo | grep -E '(Cached|Buffers|Slab)'
# OOM killer logs
dmesg | grep -i "out of memory"
sudo journalctl -xe | grep -i oom
# Recent deployments (Kubernetes)
kubectl rollout history deployment/[name]
# Recent pod restarts
kubectl get pods --all-namespaces --field-selector=status.phase=Running --sort-by=.status.startTime
# Check resource limits
kubectl describe pod [pod-name] | grep -A 5 "Limits"
# Git commits (if applicable)
git log --since="6 hours ago" --oneline
/proc/meminfo, "Available" still reasonablekubectl describe node [node]Objective: Prevent OOM, keep system stable
``bash # This is safe - kernel will rebuild cache as needed sync && echo 3 > /proc/sys/vm/drop_caches ` Expected outcome: Memory usage drops 10-20% Verification: Run free -h` again
```bash # Kubernetes kubectl rollout restart deployment/[name]
# Systemd service sudo systemctl restart [service]
# Docker container docker restart [container-id] ``` Expected outcome: Memory released, service recovers Verification: Check memory usage of new pods/processes
```bash # Kubernetes - add more replicas kubectl scale deployment [name] --replicas=[N+2]
# Cloud - add instances aws autoscaling set-desired-capacity --auto-scaling-group-name [name] --desired-capacity [N+1] `` Expected outcome: Load distributed, per-instance memory reduced Verification: kubectl top pods` shows lower memory per pod
```bash # Kubernetes - delete non-critical pods kubectl delete pod [pod-name] -n [namespace]
# Stop batch jobs kubectl delete job [job-name] ```
If above don't work: Escalate to platform team lead immediately
Objective: Stable state while investigating root cause
```yaml # Edit deployment kubectl edit deployment [name]
# Set appropriate limits resources: requests: memory: "256Mi" limits: memory: "512Mi" ```
```bash # Check current swap swapon -s
# Create swap file (if none exists) sudo fallocate -l 4G /swapfile sudo chmod 600 /swapfile sudo mkswap /swapfile sudo swapon /swapfile ``` Note: This buys time but doesn't fix the problem
```bash # Java applications - reduce heap size export JAVA_OPTS="-Xmx2g -Xms512m"
# Node.js - set max old space export NODE_OPTIONS="--max-old-space-size=2048"
# Python - limit worker memory export WORKER_MEMORY_LIMIT=512M ```
``yaml apiVersion: policy/v1 kind: PodDisruptionBudget metadata: name: [app]-pdb spec: minAvailable: 2 selector: matchLabels: app: [app] ``
If remediation makes things worse:
# Rollback deployment
kubectl rollout undo deployment/[name]
# Restore previous auto-scaling settings
kubectl apply -f autoscaling-backup.yaml
# Disable swap (if enabled as temp fix)
sudo swapoff /swapfile
sudo rm /swapfile
| Stakeholder | When | How | |-------------|------|-----| | Platform Team | Immediately | Slack #platform-oncall | | Engineering Manager | If not resolved in 30min | Slack + Phone | | Application Owners | If specific app identified | Slack team channel | | Customers | If service degraded > 15min | Status page |
Initial notification:
[INCIDENT] High Memory Usage on [node/cluster]
Impact: [service] may experience slow responses
Started: [time]
Status: Investigating memory consumption
Next update: [+15min]
Incident lead: [oncall engineer]
Slack: #incident-[ID]
# 1. Memory usage normal
free -h | grep Mem
# Available should be > 20%
# 2. No OOM events in last 10 minutes
dmesg | grep -i oom | tail -5
# 3. All pods running
kubectl get pods --all-namespaces | grep -v Running
# 4. Application metrics normal
# Check dashboard: [URL]
# 5. No pending pod evictions
kubectl get events --all-namespaces | grep Evicted
All checks passing? ✓ Incident resolved
OOMKiller: OOM events occurredHighSwapUsage: Swapping indicates memory pressurePodEvicted: Pods evicted due to resource pressureMemTotal : Total physical RAM
MemFree : Unused RAM
MemAvailable : RAM available for starting new applications (includes reclaimable cache)
Buffers : Temporary storage for block devices
Cached : Page cache from files
Slab : Kernel data structures
SwapTotal : Total swap space
SwapFree : Unused swap space
Key metric: MemAvailable (not MemFree)
# Pod memory requests vs limits
kubectl describe node [node] | grep -A 5 "Allocated resources"
# Memory pressure eviction thresholds
kubectl get nodes -o json | jq '.items[].status.allocatable.memory'
# Top memory-consuming namespaces
kubectl top pods -A | awk '{print $1, $4}' | sort -k2 -rh | head -20
| Date | Author | Change | |------|--------|--------| | 2025-10-27 | Platform Team | Initial version | | 2025-10-27 | SRE | Added Kubernetes-specific steps |
Found this runbook helpful? Found an issue?