Checkout SLOs and alerts
The three objectives the checkout path commits to, and the burn-rate alerts that page when a budget burns too fast.
Checkout is the path from POST /checkout to the reply that carries a captured payment. Three objectives cover it: the request succeeds, it succeeds fast, and the captured amount matches the order. Each objective has a 30-day window and an error budget. Alerts fire on how fast the budget burns, never on a raw error count. A quiet night does not page; a bad deploy pages within minutes.
Assumptions
Objectives
Availability and latency are measured at the gateway, so a slow payment provider counts against us. That is deliberate: the customer does not know who was slow. Capture correctness is measured in the ledger one hour after the order. Its budget burns slowly, so its alert is a ticket, not a page.
Checkout · 30 days to 12 Sep 2026
POST /checkout responses that are not 5xx or timeout, over all responses
POST /checkout responses under 1,500 ms, over all successful responses
Orders whose captured amount equals the order total within 1 hour, over all placed orders
Where the latency budget goes
The latency objective is the one we miss. The p99 of POST /checkout is 2,100 ms, and 1,900 ms of it is the capture call to the payment provider. The cart endpoints are not part of the objective; they are here to show that the checkout tail is not a platform problem. The fix is in flight: capture moves off the request path in October, and the objective stays at 1,500 ms until then.
Checkout path latency · 30 days to 12 Sep 2026
Alert rules
Every alert uses two windows. It fires only when both the long window and the short window burn faster than the threshold. The long window keeps a two-minute blip from paging; the short window clears the alert within minutes of recovery. A fast burn at 14.4× empties the 30-day budget in two days, so it pages. A trickle at 1× empties it exactly at the end of the window, so it opens a ticket for the next working day.
Burn-rate alerts on the checkout objectives
| Alert | Long window | Short window | Burn rate | Budget spent in long window | Action |
|---|---|---|---|---|---|
| Availability fast burn | 1 h | 5 min | 14.4× | 2% | Page |
| Availability slow burn | 6 h | 30 min | 6× | 5% | Page |
| Availability trickle | 3 d | 6 h | 1× | 10% | Ticket |
| Latency fast burn | 1 h | 5 min | 14.4× | 2% | Page |
| Latency slow burn | 6 h | 30 min | 6× | 5% | Ticket |
| Capture mismatch | 6 h | 1 h | 3× | 2.5% | Ticket |
Pages route to the checkout on-call in PagerDuty service CHK. Tickets land in the CHECKOUT Jira board with the alert name as the title. An alert on a paused objective, marked in the SLO dashboard, never pages.
What the on-call does with a page
A page means the budget will run out in days, not that checkout is down. The first question is whether the cause is ours. The provider status page and the deploy log answer it in under two minutes, and each answer has one move. The on-call does not tune the alert thresholds during an incident; a threshold change is a pull request to the SLO repo.