Skip to content
chiltepin

Generated from: “What happens if we lose the primary region.”

Losing the primary region disaster recovery plan

AI-generated example · 2026-09-14. The source passes chiltepin check. System details and measurements are illustrative; review them before adapting this document.

DOCUMENTRUNBOOK

Losing the primary region

What happens in the first hour after us-east-1 goes away, who decides, and what the last drill proved.

The platform runs active-passive across two regions. us-east-1 takes every write; eu-west-1 serves European reads and holds an async Postgres replica and mirrored Kafka topics. Losing the primary region means promoting the secondary, and that is a human decision. Automation prepares every step, but it never promotes on its own. A false promotion splits the write path and costs more than 20 minutes of downtime.

SECTION 01 · Note

Assumptions

Note
The request did not name the regions or the targets. This document assumes AWS us-east-1 as primary and eu-west-1 as secondary. The targets are an RTO of 30 minutes and an RPO of 5 seconds, agreed with the business in Q2 2026. Region loss means the AWS status page or three independent probes report the region unreachable for five minutes. A single failed zone is not a region loss and is handled by the zone layout, not by this runbook.

The first hour

The timeline is measured from the first probe failure. The declaration at T+5 is the step that takes judgment; every step after it is a script or a console click. The 30-minute RTO ends at T+30, when writes are accepted in eu-west-1. Full service, including the search index and the outbound webhooks, returns by T+55.

SECTION 02 · Roadmap

Failover from us-east-1 to eu-west-1

T+0
done
Probes fail
Three probes from three networks report us-east-1 unreachable; the platform on-call is paged
T+5
done
Region loss declared
The on-call and the incident lead agree on the criteria and open the DR incident
T+7
current
Writes frozen
The gateway returns 503 for POST and PUT in both regions; reads continue from eu-west-1
T+10
next
Postgres promoted
The eu-west-1 replica is promoted; the last applied LSN is recorded for the loss estimate
T+15
next
Services repointed
orders, billing, and auth in eu-west-1 restart with the promoted primary; Kafka consumers resume from the mirrored offsets
T+20
next
DNS switched
Route 53 health check flips; the 60-second TTL drains us-east-1 clients
T+30
next
Writes reopened
The gateway accepts writes in eu-west-1; the RTO clock stops here
T+55
next
Full service
The search index rebuilds from Kafka and webhooks replay from the outbox
Legenddonecurrentnext

The decisions on the way

Three decisions sit inside the first 30 minutes, and each has a wrong answer that costs more than waiting. Promoting on a partial outage splits writes. Promoting with a replica lag above the RPO loses orders that customers already paid for, so that path stops and calls the business. Switching DNS before the services are healthy sends customers to errors in a second region.

SECTION 03 · Flowchart

Decision points before writes reopen

FLOW
Flowchart: 9 stepsThree probes failfor 5 minWhole region, notone zone or oneZone runbook; nofailoverReplica lag under5 s at the freeze?Call the incidentlead and thePromote Postgres;restart services inAll services passreadiness inFix the failingservice; DNS staysSwitch DNS; reopenwrites12345678
1no2yes3yes4no5loss accepted6yes7no8retry
Legendstartstepdecision (diamond)exitnextoptionalerror pathhappy path

What we have proved

The quarterly drill runs the full failover on a Sunday at 06:00 UTC with real traffic. The last drill on 14 Jul 2026 reopened writes at T+24. Two items fail: the search index rebuild took 48 minutes instead of 25, and the webhook replay sent 212 duplicates to one partner. Both are open tickets with owners and both must pass before the October drill.

SECTION 04 · Checklist

Region-loss readiness · drill of 14 Jul 2026

DR standard v3
Data
pass
Postgres replica lag under 5 s for 30 daysGrafana panel pg-replication, max 2.4 s
pass
Kafka mirror lag under 10 s for 30 daysMirrorMaker 2 checkpoint metric, max 6 s
pass
Promotion script runs in under 3 minutesdrill log, 2 min 41 s
partial
Object storage replicated within 15 minutesS3 CRR met 15 min for 99.2% of objects; two buckets lack CRR
Services
pass
Every service has an eu-west-1 deployment at 50% of primary capacityArgo CD app list, 31 of 31
pass
Secrets and config present in eu-west-1Vault replication audit, 0 missing
fail
Search index rebuilds in under 25 minutesdrill: 48 min; ticket SRCH-412 splits the rebuild by shard
fail
Webhook replay sends no duplicatesdrill: 212 duplicates to one partner; ticket PLAT-988 adds the idempotency key
People and process
pass
Two people can declare region loss at any houron-call rota, platform plus incident lead
pass
Drill run every quarter with real traffic14 Jul 2026, next on 12 Oct 2026
pending
Failback runbook rehearsedscheduled for the October drill
7 pass · 2 fail · 1 partial · 1 pendingpass rate 68%

Targets against the last drill

The RTO holds with six minutes to spare. The RPO holds because replica lag stays well under five seconds. The number is a ceiling, not a guarantee; a burst of writes during the freeze is the case that pushes it. Failback to us-east-1 is not in the targets and takes a planned window.

SECTION 05 · Metrics

Targets and the drill of 14 Jul 2026

30 min
RTO target
24 min
RTO in the last drill
▲ −6 min
5 s
RPO target
2.4 s
Max replica lag in 30 days
▲ −2.6 s
2 of 11
Checklist items failing
▲ −1
View the Markdown
```meta
title: Losing the primary region
subtitle: What happens in the first hour after us-east-1 goes away, who decides, and what the last drill proved.
tag: RUNBOOK
```

The platform runs active-passive across two regions. `us-east-1` takes every write; `eu-west-1` serves European reads and holds an async Postgres replica and mirrored Kafka topics. Losing the primary region means promoting the secondary, and that is a human decision. Automation prepares every step, but it never promotes on its own. A false promotion splits the write path and costs more than 20 minutes of downtime.

```callout
tone: note
title: Assumptions
body: "The request did not name the regions or the targets. This document assumes AWS us-east-1 as primary and eu-west-1 as secondary. The targets are an RTO of 30 minutes and an RPO of 5 seconds, agreed with the business in Q2 2026. Region loss means the AWS status page or three independent probes report the region unreachable for five minutes. A single failed zone is not a region loss and is handled by the zone layout, not by this runbook."
```

## The first hour

The timeline is measured from the first probe failure. The declaration at T+5 is the step that takes judgment; every step after it is a script or a console click. The 30-minute RTO ends at T+30, when writes are accepted in `eu-west-1`. Full service, including the search index and the outbound webhooks, returns by T+55.

```timeline
id: region-failover
title: Failover from us-east-1 to eu-west-1
items:
  - "[done] T+0 · Probes fail · Three probes from three networks report us-east-1 unreachable; the platform on-call is paged"
  - "[done] T+5 · Region loss declared · The on-call and the incident lead agree on the criteria and open the DR incident"
  - "[current] T+7 · Writes frozen · The gateway returns 503 for POST and PUT in both regions; reads continue from eu-west-1"
  - "[next] T+10 · Postgres promoted · The eu-west-1 replica is promoted; the last applied LSN is recorded for the loss estimate"
  - "[next] T+15 · Services repointed · orders, billing, and auth in eu-west-1 restart with the promoted primary; Kafka consumers resume from the mirrored offsets"
  - "[next] T+20 · DNS switched · Route 53 health check flips; the 60-second TTL drains us-east-1 clients"
  - "[next] T+30 · Writes reopened · The gateway accepts writes in eu-west-1; the RTO clock stops here"
  - "[next] T+55 · Full service · The search index rebuilds from Kafka and webhooks replay from the outbox"
```

## The decisions on the way

Three decisions sit inside the first 30 minutes, and each has a wrong answer that costs more than waiting. Promoting on a partial outage splits writes. Promoting with a replica lag above the RPO loses orders that customers already paid for, so that path stops and calls the business. Switching DNS before the services are healthy sends customers to errors in a second region.

```flow
id: region-decisions
title: Decision points before writes reopen
dir: TB
nodes:
  - { id: probes, col: 2, row: 1, kind: start, label: "Three probes fail for 5 min" }
  - { id: whole, col: 2, row: 2, kind: decision, label: "Whole region, not one zone or one service?" }
  - { id: zone, col: 1, row: 2, kind: end, label: "Zone runbook; no failover" }
  - { id: lag, col: 2, row: 3, kind: decision, label: "Replica lag under 5 s at the freeze?" }
  - { id: business, col: 3, row: 3, kind: process, label: "Call the incident lead and the business owner; accept the loss or wait" }
  - { id: promote, col: 2, row: 4, kind: process, label: "Promote Postgres; restart services in eu-west-1" }
  - { id: healthy, col: 2, row: 5, kind: decision, label: "All services pass readiness in eu-west-1?" }
  - { id: fix, col: 3, row: 5, kind: process, label: "Fix the failing service; DNS stays put" }
  - { id: dns, col: 2, row: 6, kind: end, label: "Switch DNS; reopen writes" }
edges:
  - probes -> whole
  - whole -x-> zone: "no"
  - whole -> lag: "yes"
  - lag -> promote: "yes"
  - lag -x-> business: "no"
  - business --> promote: "loss accepted"
  - promote -> healthy
  - healthy -> dns: "yes"
  - healthy -x-> fix: "no"
  - fix --> healthy: "retry"
```

## What we have proved

The quarterly drill runs the full failover on a Sunday at 06:00 UTC with real traffic. The last drill on 14 Jul 2026 reopened writes at T+24. Two items fail: the search index rebuild took 48 minutes instead of 25, and the webhook replay sent 212 duplicates to one partner. Both are open tickets with owners and both must pass before the October drill.

```checklist
id: dr-readiness
title: Region-loss readiness · drill of 14 Jul 2026
standard: DR standard v3
groups:
  - label: Data
    items:
      - "[pass] Postgres replica lag under 5 s for 30 days — Grafana panel pg-replication, max 2.4 s"
      - "[pass] Kafka mirror lag under 10 s for 30 days — MirrorMaker 2 checkpoint metric, max 6 s"
      - "[pass] Promotion script runs in under 3 minutes — drill log, 2 min 41 s"
      - "[partial] Object storage replicated within 15 minutes — S3 CRR met 15 min for 99.2% of objects; two buckets lack CRR"
  - label: Services
    items:
      - "[pass] Every service has an eu-west-1 deployment at 50% of primary capacity — Argo CD app list, 31 of 31"
      - "[pass] Secrets and config present in eu-west-1 — Vault replication audit, 0 missing"
      - "[fail] Search index rebuilds in under 25 minutes — drill: 48 min; ticket SRCH-412 splits the rebuild by shard"
      - "[fail] Webhook replay sends no duplicates — drill: 212 duplicates to one partner; ticket PLAT-988 adds the idempotency key"
  - label: People and process
    items:
      - "[pass] Two people can declare region loss at any hour — on-call rota, platform plus incident lead"
      - "[pass] Drill run every quarter with real traffic — 14 Jul 2026, next on 12 Oct 2026"
      - "[pending] Failback runbook rehearsed — scheduled for the October drill"
```

## Targets against the last drill

The RTO holds with six minutes to spare. The RPO holds because replica lag stays well under five seconds. The number is a ceiling, not a guarantee; a burst of writes during the freeze is the case that pushes it. Failback to `us-east-1` is not in the targets and takes a planned window.

```stats
id: dr-targets
title: Targets and the drill of 14 Jul 2026
stats:
  - { value: "30 min", label: RTO target }
  - { value: "24 min", label: RTO in the last drill, delta: "−6 min", trend: up, accent: green }
  - { value: "5 s", label: RPO target }
  - { value: "2.4 s", label: Max replica lag in 30 days, delta: "−2.6 s", trend: up, accent: green }
  - { value: "2 of 11", label: Checklist items failing, delta: "−1", trend: up, accent: amber }
```