Checkout outage — 14 Aug
INC-2026-0814 · SEV-1 · Orders API unavailable for 47 minutes
On 14 Aug the Orders API returned errors for 47 minutes and no customer could complete checkout. The trigger was a routine deploy that enabled the price-check-v2 flag. The flag made Orders call Pricing synchronously on every request. Pricing ran out of database connections, and retries without backoff turned a slowdown into a full outage.
Assumptions in this document: all times are UTC. The deploy, flag, and service names come from the incident channel. The customer counts come from the checkout funnel dashboard, not from support tickets.
Detection to resolution
Fourteen minutes passed between the first symptom and the page. The pool wait alert existed but routed to a Slack channel, not to the pager, so the first human signal was the customer-facing error rate. The diagnosis itself took 14 minutes. The deploy shipped four unrelated changes, and the flag flip was not in the release notes.
How the failure spread
Each checkout held a Pricing connection open for up to 20 seconds and retried three times without backoff, so one slow query cost four connections. With the pool at 20 connections per pod, 45 concurrent checkouts were enough to block every Pricing pod. The load balancer then removed Orders pods whose threads were stuck in retries, which pushed the remaining pods over the same cliff.
Why it broke
The direct cause was the flag flip, but the outage needed all four bones. Without the retry storm, the pool degrades gracefully. Without the shared thread pool, the load balancer keeps healthy pods. With a paged alert, the response starts nine minutes earlier.
Impact
Failed checkouts are counted once per session. The revenue figure uses the average order value of the prior seven days. It does not subtract orders that customers retried after 10:28. The funnel shows that 41% of affected sessions converted later that day.
Remediation
- 1Make the Pricing call asynchronous
Orders reads a cached price and enqueues a verification job; a stale price never blocks checkout. Owner M. Okafor, due 28 Aug.
- 2Add backoff, jitter, and a circuit breaker to the Pricing client
Retries use exponential backoff from 200 ms with jitter, and the breaker opens after 50% failures in 30 s. Owner R. Lindqvist, due 21 Aug.
yamlretry: { attempts: 3, backoff: exponential, base: 200ms, jitter: full } - 3Route the pool wait alert to the pager
Pool wait over 500 ms for 2 minutes pages the Pricing on-call. Owner A. Ferreira, done 15 Aug.
- 4Separate the health check thread pool
The load balancer probe answers from its own pool so busy request threads never mark a pod unhealthy. Owner R. Lindqvist, due 4 Sep.
- 5Require a canary stage for every flag above 10%
The flag service refuses a jump from under 10% to over 50% without a 30 minute canary. Policy shipped 15 Aug; enforcement owner A. Ferreira, due 11 Sep.
Applies to all services, not only Orders.
Items 1 and 2 remove the mechanism; items 3 and 4 shorten detection and stop the cascade; item 5 removes the trigger class. The team accepts that item 1 can show a price up to 60 seconds stale at checkout, which is within the pricing contract.