Skip to content
chiltepin

Generated from: “Define SLOs for checkout and how we alert on them.”

Checkout SLOs and alerts

AI-generated example · 2026-09-14. The source passes chiltepin check. System details and measurements are illustrative; review them before adapting this document.

DOCUMENTRUNBOOK

Checkout SLOs and alerts

The three objectives the checkout path commits to, and the burn-rate alerts that page when a budget burns too fast.

Checkout is the path from POST /checkout to the reply that carries a captured payment. Three objectives cover it: the request succeeds, it succeeds fast, and the captured amount matches the order. Each objective has a 30-day window and an error budget. Alerts fire on how fast the budget burns, never on a raw error count. A quiet night does not page; a bad deploy pages within minutes.

SECTION 01 · Note

Assumptions

Note
The request named no numbers, so this document sets them. Checkout runs at about 1,200 completed checkouts per minute at peak and 180 at night. Availability and latency come from the gateway access log for POST /checkout. Capture correctness comes from the payments ledger. Windows are 30 days rolling. The targets are the ones the payments and storefront teams agreed for Q3 2026; the current values are the 30 days to 12 Sep 2026.

Objectives

Availability and latency are measured at the gateway, so a slow payment provider counts against us. That is deliberate: the customer does not know who was slow. Capture correctness is measured in the ledger one hour after the order. Its budget burns slowly, so its alert is a ticket, not a page.

SECTION 02 · Service objectives

Checkout · 30 days to 12 Sep 2026

Availability30d

POST /checkout responses that are not 5xx or timeout, over all responses

Target99.9%
Current99.94%
58% of error budget used
Latency30d

POST /checkout responses under 1,500 ms, over all successful responses

Target99%
Current98.7%
error budget exhausted
Capture correctness30d

Orders whose captured amount equals the order total within 1 hour, over all placed orders

Target99.99%
Current99.996%
40% of error budget used

Where the latency budget goes

The latency objective is the one we miss. The p99 of POST /checkout is 2,100 ms, and 1,900 ms of it is the capture call to the payment provider. The cart endpoints are not part of the objective; they are here to show that the checkout tail is not a platform problem. The fix is in flight: capture moves off the request path in October, and the objective stays at 1,500 ms until then.

SECTION 03 · Latency percentiles

Checkout path latency · 30 days to 12 Sep 2026

PERCENTILES
Latency percentiles: 4 rows10ms20ms50ms100ms200ms500ms1000ms2000ms5000msPOST /checkoutSLO 1500ms4202100POST /payments/capture (provider hop)SLO 1500ms3101900POST /cart/items38260GET /cart22150
Legendp50p90 · p95p99p99 over the SLOmaxSLO

Alert rules

Every alert uses two windows. It fires only when both the long window and the short window burn faster than the threshold. The long window keeps a two-minute blip from paging; the short window clears the alert within minutes of recovery. A fast burn at 14.4× empties the 30-day budget in two days, so it pages. A trickle at 1× empties it exactly at the end of the window, so it opens a ticket for the next working day.

SECTION 04 · Comparison

Burn-rate alerts on the checkout objectives

AlertLong windowShort windowBurn rateBudget spent in long windowAction
Availability fast burn1 h5 min14.4×2%Page
Availability slow burn6 h30 min6×5%Page
Availability trickle3 d6 h1×10%Ticket
Latency fast burn1 h5 min14.4×2%Page
Latency slow burn6 h30 min6×5%Ticket
Capture mismatch6 h1 h3×2.5%Ticket

Pages route to the checkout on-call in PagerDuty service CHK. Tickets land in the CHECKOUT Jira board with the alert name as the title. An alert on a paused objective, marked in the SLO dashboard, never pages.

What the on-call does with a page

A page means the budget will run out in days, not that checkout is down. The first question is whether the cause is ours. The provider status page and the deploy log answer it in under two minutes, and each answer has one move. The on-call does not tune the alert thresholds during an incident; a threshold change is a pull request to the SLO repo.

SECTION 05 · Flowchart

From page to first action

FLOW
Flowchart: 7 stepsPage: fast or slowburnProvider statuspage degraded?Route capture to thefallback providerCheckout deploy inthe last 60 min?Abort the rollout;stable versionOpen an incident;page the paymentsWatch the shortwindow; the alert12345
1yes2no3yes4no5after mitigation
Legendstartstepdecision (diamond)nexterror pathhappy path
View the Markdown
```meta
title: Checkout SLOs and alerts
subtitle: The three objectives the checkout path commits to, and the burn-rate alerts that page when a budget burns too fast.
tag: RUNBOOK
```

Checkout is the path from `POST /checkout` to the reply that carries a captured payment. Three objectives cover it: the request succeeds, it succeeds fast, and the captured amount matches the order. Each objective has a 30-day window and an error budget. Alerts fire on how fast the budget burns, never on a raw error count. A quiet night does not page; a bad deploy pages within minutes.

```callout
tone: note
title: Assumptions
body: "The request named no numbers, so this document sets them. Checkout runs at about 1,200 completed checkouts per minute at peak and 180 at night. Availability and latency come from the gateway access log for POST /checkout. Capture correctness comes from the payments ledger. Windows are 30 days rolling. The targets are the ones the payments and storefront teams agreed for Q3 2026; the current values are the 30 days to 12 Sep 2026."
```

## Objectives

Availability and latency are measured at the gateway, so a slow payment provider counts against us. That is deliberate: the customer does not know who was slow. Capture correctness is measured in the ledger one hour after the order. Its budget burns slowly, so its alert is a ticket, not a page.

```slo
id: checkout-objectives
title: Checkout · 30 days to 12 Sep 2026
items:
  - { name: Availability, sli: "POST /checkout responses that are not 5xx or timeout, over all responses", target: 99.9%, current: 99.94%, window: 30d, budget: 0.58 }
  - { name: Latency, sli: "POST /checkout responses under 1,500 ms, over all successful responses", target: 99%, current: 98.7%, window: 30d, budget: 1 }
  - { name: Capture correctness, sli: "Orders whose captured amount equals the order total within 1 hour, over all placed orders", target: 99.99%, current: 99.996%, window: 30d, budget: 0.4 }
```

## Where the latency budget goes

The latency objective is the one we miss. The p99 of `POST /checkout` is 2,100 ms, and 1,900 ms of it is the capture call to the payment provider. The cart endpoints are not part of the objective; they are here to show that the checkout tail is not a platform problem. The fix is in flight: capture moves off the request path in October, and the objective stays at 1,500 ms until then.

```percentiles
id: checkout-tail
title: Checkout path latency · 30 days to 12 Sep 2026
unit: ms
scale: log
rows:
  - { label: POST /checkout, p50: 420, p90: 900, p95: 1200, p99: 2100, max: 9800, slo: 1500 }
  - { label: "POST /payments/capture (provider hop)", p50: 310, p90: 720, p95: 980, p99: 1900, max: 9200, slo: 1500, accent: red }
  - { label: POST /cart/items, p50: 38, p90: 85, p95: 120, p99: 260, max: 2100, accent: teal }
  - { label: GET /cart, p50: 22, p90: 48, p95: 70, p99: 150, max: 1400, accent: teal }
```

## Alert rules

Every alert uses two windows. It fires only when both the long window and the short window burn faster than the threshold. The long window keeps a two-minute blip from paging; the short window clears the alert within minutes of recovery. A fast burn at 14.4× empties the 30-day budget in two days, so it pages. A trickle at 1× empties it exactly at the end of the window, so it opens a ticket for the next working day.

```table
id: checkout-alerts
title: Burn-rate alerts on the checkout objectives
columns: [Alert, Long window, Short window, Burn rate, Budget spent in long window, Action]
rows:
  - [Availability fast burn, 1 h, 5 min, "14.4×", "2%", { v: Page, tone: neg }]
  - [Availability slow burn, 6 h, 30 min, "6×", "5%", { v: Page, tone: neg }]
  - [Availability trickle, 3 d, 6 h, "1×", "10%", { v: Ticket, tone: warn }]
  - [Latency fast burn, 1 h, 5 min, "14.4×", "2%", { v: Page, tone: neg }]
  - [Latency slow burn, 6 h, 30 min, "6×", "5%", { v: Ticket, tone: warn }]
  - [Capture mismatch, 6 h, 1 h, "3×", "2.5%", { v: Ticket, tone: warn }]
note: "Pages route to the checkout on-call in PagerDuty service CHK. Tickets land in the CHECKOUT Jira board with the alert name as the title. An alert on a paused objective, marked in the SLO dashboard, never pages."
```

## What the on-call does with a page

A page means the budget will run out in days, not that checkout is down. The first question is whether the cause is ours. The provider status page and the deploy log answer it in under two minutes, and each answer has one move. The on-call does not tune the alert thresholds during an incident; a threshold change is a pull request to the SLO repo.

```flow
id: checkout-page-response
title: From page to first action
dir: LR
nodes:
  - { id: page, col: 1, row: 1, kind: start, label: "Page: fast or slow burn" }
  - { id: provider, col: 2, row: 1, kind: decision, label: "Provider status page degraded?" }
  - { id: fallback, col: 3, row: 1, kind: process, label: "Route capture to the fallback provider" }
  - { id: deploy, col: 3, row: 2, kind: decision, label: "Checkout deploy in the last 60 min?" }
  - { id: rollback, col: 4, row: 2, kind: process, label: "Abort the rollout; stable version serves" }
  - { id: incident, col: 4, row: 3, kind: process, label: "Open an incident; page the payments on-call" }
  - { id: watch, col: 5, row: 1, kind: end, label: "Watch the short window; the alert clears under 1×" }
edges:
  - page -> provider
  - provider -> fallback: "yes"
  - provider -> deploy: "no"
  - deploy -> rollback: "yes"
  - deploy -x-> incident: "no"
  - fallback -> watch
  - rollback -> watch
  - incident -> watch: "after mitigation"
```