Skip to content
chiltepin

Generated from: “What we need for Black Friday at five times normal traffic.”

Black Friday capacity at 5× normal traffic

AI-generated example · 2026-09-14. The source passes chiltepin check. System details and measurements are illustrative; review them before adapting this document.

DOCUMENTPLAN

Black Friday capacity at 5× normal traffic

What the storefront needs to hold five times normal peak on 27 Nov 2026, and the plan that gets it there.

Normal peak is 18,000 requests per second at the edge and 1,200 completed checkouts per minute. Black Friday planning assumes five times that: 90,000 requests per second and 6,000 checkouts per minute, sustained for four hours from 09:00 to 13:00 UTC. Every number below derives from those two, plus 50% headroom so that one lost availability zone does not become the incident.

SECTION 01 · Note

Assumptions

Note
Normal peak is the 95th percentile of the daily peak over the 30 days to 12 Sep 2026. Five times is the multiplier the commerce team set from last year's 3.8× and this year's marketing spend. The edge cache absorbs 65% of requests at normal peak and we assume the same hit rate; promotional pages are pre-warmed. One orders-api pod serves 500 requests per second at 70% CPU in the load test of 2 Sep 2026. The payment provider's contracted limit is 150 captures per second.

The numbers

The bottom line is 96 orders-api pods and 1,800 Postgres write IOPS. The payment provider is the tightest constraint. 6,000 checkouts per minute is 100 captures per second, two thirds of the contracted limit, and a retry storm doubles it. We raise the limit rather than trust the margin.

SECTION 02 · Capacity math

Capacity at 5× normal peak

Normal peak at the edge
18,000 rps
Normal peak checkouts
1,200 / min
Black Friday multiplier
5×
Edge cache hit rate
65%
One orders-api pod
500 rps
Writes per checkout
12 (order, items, payment, outbox)
Headroom
1.5×
Edge requests18,000 × 590,000 rps
Requests that reach orders-api90,000 × (1 − 0.65)31,500 rps
orders-api pods at 5×31,500 / 50063 pods
Checkouts1,200 × 56,000 / min ≈ 100 / s
Postgres write IOPS100 × 121,200 IOPS
Captures at the provider100 / s, retries ×2 worst case200 / s
Provision for
96 orders-api pods · 1,800 write IOPS · 300 captures / s at the provider

Where we stand today

Each tier is drawn as a multiple of normal peak. The edge and Redis already clear the 7.5× line. orders-api is capped by the cluster's node group, not by the pods. Postgres and the payment provider are the two tiers that need a contract or an instance change. Both have a lead time longer than a sprint.

SECTION 03 · Chart

Capacity by tier as a multiple of normal peak

CHART
Chart0× normal peak2.5× normal peak5× normal peak7.5× normal peak10× normal peak12.5× normal peakEdge cacheorders-apiPostgres writesRedisKafkaPayment provider12× normal peak2.5× normal peak3× normal peak9× normal peak6× normal peak1.5× normal peak7.5× normal peak7.5× normal peak7.5× normal peak7.5× normal peak7.5× normal peak7.5× normal peak
LegendCapacity todayNeeded on Black Friday

What changes, who owns it, and when it is ready

Every change lands by week 46 so the 5× load test in week 47 runs against the final shape. The freeze starts on 23 Nov. After that, only a rollback ships.

SECTION 04 · Comparison

Changes per tier

TierTodayBlack FridayChangeOwnerReady by
orders-api12 pods, node group max 4096 pods, node group max 120Raise the node group ceiling; pre-scale to 64 pods at 08:00Platform6 Nov 2026
Postgresdb.r6g.2xlarge, 3,000 IOPSdb.r6g.4xlarge, 6,000 IOPSResize the primary in the 1 Nov maintenance window; replica firstData1 Nov 2026
Payment provider150 captures / s300 captures / sContract amendment; signed limit in writingPayments30 Oct 2026
Kafka6 partitions per topic12 partitions per topicRepartition orders.* topics; consumers scale to 12Platform13 Nov 2026
Edge cache65% hit rate70% hit ratePre-warm the top 2,000 product pages at 07:00Storefront20 Nov 2026
Redis2 × cache.m6g.largeNo changeClears 9× todayPlatform—

Ready by is the date the change is live in production and covered by a load test, not the date the ticket closes.

The plan by week

Two load tests bound the plan. The 3× test in week 44 proves the resized Postgres and the new partitions. The 5× test in week 47 is the gate for the freeze; a fail there means a smaller promotion, and marketing knows that.

SECTION 05 · Schedule

Black Friday readiness · weeks 43 to 48

GANTT
ScheduleWk 43Wk 44Wk 45Wk 46Wk 47Wk 48Provider limitraisedPostgres resizeLoad test at 3×Node group ceilingand pre-scaleKafka repartitionEdge pre-warm jobLoad test at 5×Freeze and BlackFriday
Legendplannedin progressmilestone

What can still go wrong

The register holds the three risks that no line in the plan removes. Each has an owner and the move we make if it lands.

SECTION 06 · Risk register

Open risks

highThe provider raises the limit on paper but throttles at 200 / sPaymentsmitigating
L: med · I: high

Mitigation: Run 250 captures / s against the provider sandbox in week 46; keep the fallback provider warm at 20% of traffic.

mediumTraffic exceeds 5× in the first 20 minutesPlatformmitigating
L: med · I: med

Mitigation: Queue checkout at the gateway above 7,000 / min; the customer sees a wait page, never an error.

highPostgres resize slips past 1 NovDataopen
L: low · I: high

Mitigation: Second window on 8 Nov; after that the 3× test runs on the current instance and the plan drops to 4×.

View the Markdown
```meta
title: Black Friday capacity at 5× normal traffic
subtitle: What the storefront needs to hold five times normal peak on 27 Nov 2026, and the plan that gets it there.
tag: PLAN
```

Normal peak is 18,000 requests per second at the edge and 1,200 completed checkouts per minute. Black Friday planning assumes five times that: 90,000 requests per second and 6,000 checkouts per minute, sustained for four hours from 09:00 to 13:00 UTC. Every number below derives from those two, plus 50% headroom so that one lost availability zone does not become the incident.

```callout
tone: note
title: Assumptions
body: "Normal peak is the 95th percentile of the daily peak over the 30 days to 12 Sep 2026. Five times is the multiplier the commerce team set from last year's 3.8× and this year's marketing spend. The edge cache absorbs 65% of requests at normal peak and we assume the same hit rate; promotional pages are pre-warmed. One orders-api pod serves 500 requests per second at 70% CPU in the load test of 2 Sep 2026. The payment provider's contracted limit is 150 captures per second."
```

## The numbers

The bottom line is 96 orders-api pods and 1,800 Postgres write IOPS. The payment provider is the tightest constraint. 6,000 checkouts per minute is 100 captures per second, two thirds of the contracted limit, and a retry storm doubles it. We raise the limit rather than trust the margin.

```envelope
id: bf-math
title: Capacity at 5× normal peak
assumptions:
  - { label: Normal peak at the edge, value: "18,000 rps" }
  - { label: Normal peak checkouts, value: "1,200 / min" }
  - { label: Black Friday multiplier, value: "5×" }
  - { label: Edge cache hit rate, value: "65%" }
  - { label: One orders-api pod, value: "500 rps" }
  - { label: Writes per checkout, value: "12 (order, items, payment, outbox)" }
  - { label: Headroom, value: "1.5×" }
steps:
  - { label: Edge requests, calc: "18,000 × 5", result: "90,000 rps" }
  - { label: Requests that reach orders-api, calc: "90,000 × (1 − 0.65)", result: "31,500 rps" }
  - { label: orders-api pods at 5×, calc: "31,500 / 500", result: "63 pods" }
  - { label: Checkouts, calc: "1,200 × 5", result: "6,000 / min ≈ 100 / s" }
  - { label: Postgres write IOPS, calc: "100 × 12", result: "1,200 IOPS" }
  - { label: Captures at the provider, calc: "100 / s, retries ×2 worst case", result: "200 / s" }
result: { label: Provision for, value: "96 orders-api pods · 1,800 write IOPS · 300 captures / s at the provider" }
```

## Where we stand today

Each tier is drawn as a multiple of normal peak. The edge and Redis already clear the 7.5× line. orders-api is capped by the cluster's node group, not by the pods. Postgres and the payment provider are the two tiers that need a contract or an instance change. Both have a lead time longer than a sprint.

```chart
id: bf-headroom
title: Capacity by tier as a multiple of normal peak
kind: bar
unit: "× normal peak"
labels: [Edge cache, orders-api, Postgres writes, Redis, Kafka, Payment provider]
series:
  - { label: Capacity today, accent: gray, values: [12, 2.5, 3, 9, 6, 1.5] }
  - { label: Needed on Black Friday, accent: navy, values: [7.5, 7.5, 7.5, 7.5, 7.5, 7.5] }
```

## What changes, who owns it, and when it is ready

Every change lands by week 46 so the 5× load test in week 47 runs against the final shape. The freeze starts on 23 Nov. After that, only a rollback ships.

```table
id: bf-changes
title: Changes per tier
columns: [Tier, Today, Black Friday, Change, Owner, Ready by]
rows:
  - [orders-api, "12 pods, node group max 40", "96 pods, node group max 120", "Raise the node group ceiling; pre-scale to 64 pods at 08:00", Platform, 6 Nov 2026]
  - [Postgres, "db.r6g.2xlarge, 3,000 IOPS", "db.r6g.4xlarge, 6,000 IOPS", "Resize the primary in the 1 Nov maintenance window; replica first", Data, 1 Nov 2026]
  - [Payment provider, "150 captures / s", "300 captures / s", "Contract amendment; signed limit in writing", Payments, 30 Oct 2026]
  - [Kafka, "6 partitions per topic", "12 partitions per topic", "Repartition orders.* topics; consumers scale to 12", Platform, 13 Nov 2026]
  - [Edge cache, "65% hit rate", "70% hit rate", "Pre-warm the top 2,000 product pages at 07:00", Storefront, 20 Nov 2026]
  - [Redis, "2 × cache.m6g.large", "No change", "Clears 9× today", Platform, "—"]
note: "Ready by is the date the change is live in production and covered by a load test, not the date the ticket closes."
```

## The plan by week

Two load tests bound the plan. The 3× test in week 44 proves the resized Postgres and the new partitions. The 5× test in week 47 is the gate for the freeze; a fail there means a smaller promotion, and marketing knows that.

```gantt
id: bf-plan
title: Black Friday readiness · weeks 43 to 48
periods: [Wk 43, Wk 44, Wk 45, Wk 46, Wk 47, Wk 48]
tasks:
  - { label: Provider limit raised, start: 0, span: 2, kind: active }
  - { label: Postgres resize, start: 0, span: 2, kind: active }
  - { label: Load test at 3×, start: 1, span: 1 }
  - { label: Node group ceiling and pre-scale, start: 2, span: 1 }
  - { label: Kafka repartition, start: 2, span: 2 }
  - { label: Edge pre-warm job, start: 3, span: 1 }
  - { label: Load test at 5×, start: 4, span: 1, kind: milestone }
  - { label: Freeze and Black Friday, start: 5, span: 1, kind: milestone }
```

## What can still go wrong

The register holds the three risks that no line in the plan removes. Each has an owner and the move we make if it lands.

```risk
id: bf-risks
title: Open risks
items:
  - { risk: "The provider raises the limit on paper but throttles at 200 / s", likelihood: med, impact: high, mitigation: "Run 250 captures / s against the provider sandbox in week 46; keep the fallback provider warm at 20% of traffic.", owner: Payments, status: mitigating }
  - { risk: "Traffic exceeds 5× in the first 20 minutes", likelihood: med, impact: med, mitigation: "Queue checkout at the gateway above 7,000 / min; the customer sees a wait page, never an error.", owner: Platform, status: mitigating }
  - { risk: "Postgres resize slips past 1 Nov", likelihood: low, impact: high, mitigation: "Second window on 8 Nov; after that the 3× test runs on the current instance and the plan drops to 4×.", owner: Data, status: open }
```