The goal is simple: lose any one region and keep serving traffic with no manual scramble. Getting there takes rehearsal, not just architecture.
Health, not heartbeat
We route on real readiness — can the region actually serve a request — rather than a simple ping. Unhealthy regions drain automatically.
- Automated traffic draining with connection-aware shifting
- Replicated state with a known, bounded lag budget
- A documented, scripted promotion path for the data tier
Rehearse on a schedule
Once a month we take a region out in production during business hours. If it is not routine, it is not reliable.