AI-generated example · 2026-09-14. The source passes chiltepin check. System details and measurements are illustrative; review them before adapting this document.
View the Markdown
```meta
title: Losing the primary region
subtitle: What happens in the first hour after us-east-1 goes away, who decides, and what the last drill proved.
tag: RUNBOOK
```
The platform runs active-passive across two regions. `us-east-1` takes every write; `eu-west-1` serves European reads and holds an async Postgres replica and mirrored Kafka topics. Losing the primary region means promoting the secondary, and that is a human decision. Automation prepares every step, but it never promotes on its own. A false promotion splits the write path and costs more than 20 minutes of downtime.
```callout
tone: note
title: Assumptions
body: "The request did not name the regions or the targets. This document assumes AWS us-east-1 as primary and eu-west-1 as secondary. The targets are an RTO of 30 minutes and an RPO of 5 seconds, agreed with the business in Q2 2026. Region loss means the AWS status page or three independent probes report the region unreachable for five minutes. A single failed zone is not a region loss and is handled by the zone layout, not by this runbook."
```
## The first hour
The timeline is measured from the first probe failure. The declaration at T+5 is the step that takes judgment; every step after it is a script or a console click. The 30-minute RTO ends at T+30, when writes are accepted in `eu-west-1`. Full service, including the search index and the outbound webhooks, returns by T+55.
```timeline
id: region-failover
title: Failover from us-east-1 to eu-west-1
items:
- "[done] T+0 · Probes fail · Three probes from three networks report us-east-1 unreachable; the platform on-call is paged"
- "[done] T+5 · Region loss declared · The on-call and the incident lead agree on the criteria and open the DR incident"
- "[current] T+7 · Writes frozen · The gateway returns 503 for POST and PUT in both regions; reads continue from eu-west-1"
- "[next] T+10 · Postgres promoted · The eu-west-1 replica is promoted; the last applied LSN is recorded for the loss estimate"
- "[next] T+15 · Services repointed · orders, billing, and auth in eu-west-1 restart with the promoted primary; Kafka consumers resume from the mirrored offsets"
- "[next] T+20 · DNS switched · Route 53 health check flips; the 60-second TTL drains us-east-1 clients"
- "[next] T+30 · Writes reopened · The gateway accepts writes in eu-west-1; the RTO clock stops here"
- "[next] T+55 · Full service · The search index rebuilds from Kafka and webhooks replay from the outbox"
```
## The decisions on the way
Three decisions sit inside the first 30 minutes, and each has a wrong answer that costs more than waiting. Promoting on a partial outage splits writes. Promoting with a replica lag above the RPO loses orders that customers already paid for, so that path stops and calls the business. Switching DNS before the services are healthy sends customers to errors in a second region.
```flow
id: region-decisions
title: Decision points before writes reopen
dir: TB
nodes:
- { id: probes, col: 2, row: 1, kind: start, label: "Three probes fail for 5 min" }
- { id: whole, col: 2, row: 2, kind: decision, label: "Whole region, not one zone or one service?" }
- { id: zone, col: 1, row: 2, kind: end, label: "Zone runbook; no failover" }
- { id: lag, col: 2, row: 3, kind: decision, label: "Replica lag under 5 s at the freeze?" }
- { id: business, col: 3, row: 3, kind: process, label: "Call the incident lead and the business owner; accept the loss or wait" }
- { id: promote, col: 2, row: 4, kind: process, label: "Promote Postgres; restart services in eu-west-1" }
- { id: healthy, col: 2, row: 5, kind: decision, label: "All services pass readiness in eu-west-1?" }
- { id: fix, col: 3, row: 5, kind: process, label: "Fix the failing service; DNS stays put" }
- { id: dns, col: 2, row: 6, kind: end, label: "Switch DNS; reopen writes" }
edges:
- probes -> whole
- whole -x-> zone: "no"
- whole -> lag: "yes"
- lag -> promote: "yes"
- lag -x-> business: "no"
- business --> promote: "loss accepted"
- promote -> healthy
- healthy -> dns: "yes"
- healthy -x-> fix: "no"
- fix --> healthy: "retry"
```
## What we have proved
The quarterly drill runs the full failover on a Sunday at 06:00 UTC with real traffic. The last drill on 14 Jul 2026 reopened writes at T+24. Two items fail: the search index rebuild took 48 minutes instead of 25, and the webhook replay sent 212 duplicates to one partner. Both are open tickets with owners and both must pass before the October drill.
```checklist
id: dr-readiness
title: Region-loss readiness · drill of 14 Jul 2026
standard: DR standard v3
groups:
- label: Data
items:
- "[pass] Postgres replica lag under 5 s for 30 days — Grafana panel pg-replication, max 2.4 s"
- "[pass] Kafka mirror lag under 10 s for 30 days — MirrorMaker 2 checkpoint metric, max 6 s"
- "[pass] Promotion script runs in under 3 minutes — drill log, 2 min 41 s"
- "[partial] Object storage replicated within 15 minutes — S3 CRR met 15 min for 99.2% of objects; two buckets lack CRR"
- label: Services
items:
- "[pass] Every service has an eu-west-1 deployment at 50% of primary capacity — Argo CD app list, 31 of 31"
- "[pass] Secrets and config present in eu-west-1 — Vault replication audit, 0 missing"
- "[fail] Search index rebuilds in under 25 minutes — drill: 48 min; ticket SRCH-412 splits the rebuild by shard"
- "[fail] Webhook replay sends no duplicates — drill: 212 duplicates to one partner; ticket PLAT-988 adds the idempotency key"
- label: People and process
items:
- "[pass] Two people can declare region loss at any hour — on-call rota, platform plus incident lead"
- "[pass] Drill run every quarter with real traffic — 14 Jul 2026, next on 12 Oct 2026"
- "[pending] Failback runbook rehearsed — scheduled for the October drill"
```
## Targets against the last drill
The RTO holds with six minutes to spare. The RPO holds because replica lag stays well under five seconds. The number is a ceiling, not a guarantee; a burst of writes during the freeze is the case that pushes it. Failback to `us-east-1` is not in the targets and takes a planned window.
```stats
id: dr-targets
title: Targets and the drill of 14 Jul 2026
stats:
- { value: "30 min", label: RTO target }
- { value: "24 min", label: RTO in the last drill, delta: "−6 min", trend: up, accent: green }
- { value: "5 s", label: RPO target }
- { value: "2.4 s", label: Max replica lag in 30 days, delta: "−2.6 s", trend: up, accent: green }
- { value: "2 of 11", label: Checklist items failing, delta: "−1", trend: up, accent: amber }
```