Generated from: “Decide between Kafka and SQS for order events; write the decision record.”
AI-generated example · 2026-09-14. The source passes chiltepin check. System details and measurements are illustrative; review them before adapting this document.
View the Markdown
```meta
title: ADR-021 — Kafka vs SQS for order events
subtitle: Which broker carries order.placed, and what the choice costs in ops and replay.
tag: ADR · Accepted
```
Orders publishes `order.placed` to three consumers today: Billing, Fulfilment, and Notifications. Fraud and Analytics join in Q4, and Fraud needs to replay 90 days of events to train its model. Volume is 40k orders a day, with a peak of 12 per second on campaign days. The current path is one SQS queue per consumer, and Orders writes each event three times. A fourth consumer means a fourth write in Orders code, and a missed write is silent.
Assumptions in this record: the platform runs on AWS. The team has no Kafka operator today. Every consumer already dedupes on `event_id`, so at-least-once delivery is acceptable on either broker.
## Options
```options
id: adr021-options
items:
- kicker: Option 1
title: SNS fan-out to SQS
how: Orders publishes once to an SNS topic; one SQS queue per consumer subscribes. FIFO topics for ordering.
pros: [No new infrastructure to run, "Consumers keep their SQS client code", "FIFO gives per-order ordering by message group"]
cons: ["14-day retention cap, so a 90-day replay needs a second copy in S3", "FIFO throughput is 300 msg/s per group without batching", "No consumer offsets: a replay means a re-publish from the archive"]
verdict: "VIABLE — fallback if MSK Serverless pricing changes"
tone: viable
- kicker: Option 2
title: Kafka on MSK Serverless
how: One topic, order-events, keyed by order_id. Each consumer runs its own consumer group.
pros: ["Replay is a consumer group reset, no archive needed", "Retention set to 90 days by config", "Per-key ordering at any throughput", "Schema registry enforces the event contract"]
cons: ["New client library in five services", "Consumer lag is a new alert to own", "About $210 a month at our volume, vs $40 for SNS + SQS"]
verdict: "CHOSEN"
tone: chosen
- kicker: Option 3
title: Self-managed Kafka on EC2
how: Three brokers across three zones, run by the platform team.
pros: [Lowest per-message cost at scale, Full control of broker config]
cons: ["Nobody on the team has run Kafka in production", "Broker upgrades and rebalances become our pages", "Break-even with MSK is above 2M events a day, 50× our volume"]
verdict: "REJECTED — the ops cost exceeds the platform team's budget"
tone: rejected
```
## Decision
```callout
tone: note
title: Decision
body: "Order events move to one Kafka topic, order-events, on MSK Serverless. Orders publishes once through the outbox relay. Each consumer owns a consumer group and its own offsets."
```
## What we accepted
```proscons
id: adr021-consequences
prosLabel: We gain
consLabel: We pay
pros:
- Orders publishes once; a new consumer is a deploy in that team's repo
- Fraud replays 90 days with a group reset in under an hour
- Ordering per order_id holds at any throughput
- The schema registry rejects a breaking payload change at publish time
cons:
- Five services learn a new client and its rebalance behaviour
- Consumer lag needs a dashboard and a page per group
- "$170 a month more than the SNS option"
- A consumer that falls 90 days behind loses events, where SQS would hold them for 14 days only
```
The trade we made explicit: $2,040 a year buys replay and a single publish path. One missed SQS write in the current design cost a two-day reconciliation in July. The team prefers a lag alert it can see to a missing event it cannot.
## The chosen topology
```block
id: adr021-topology
preset: event
groups:
- { id: consumers, col: 3, row: 1, cols: 1, rows: 5, label: Consumer groups }
nodes:
- { id: orders, col: 1, row: 3, kind: producer, name: Orders, tech: outbox relay }
- { id: topic, col: 2, row: 3, kind: topic, name: order-events, tech: "MSK Serverless · 90 d" }
- { id: billing, col: 3, row: 1, kind: consumer, name: Billing }
- { id: fulfilment, col: 3, row: 2, kind: consumer, name: Fulfilment }
- { id: notifications, col: 3, row: 3, kind: consumer, name: Notifications }
- { id: fraud, col: 3, row: 4, kind: consumer, name: Fraud, tech: Q4 }
- { id: analytics, col: 3, row: 5, kind: consumer, name: Analytics, tech: Q4 }
edges:
- orders -> topic: publish order.placed
- topic --> billing: charge the card
- topic --> fulfilment: create the shipment
- topic --> notifications: send the confirmation
- topic --> fraud: score and replay
- topic --> analytics: record the sale
```
The topic is the only thing the two sides share. Orders holds no list of consumers and no per-consumer retry logic. A consumer that is down for a week reads every event it missed when it returns.
## Follow-up work
```statustable
id: adr021-followups
columns: [Task, Update]
statuses:
- { label: done, color: success }
- { label: in progress, color: blue }
- { label: not started, color: neutral }
rows:
- { cells: [Provision MSK Serverless and the order-events topic, "Terraform merged; 90-day retention set"], status: done }
- { cells: [Register order.placed v1 in the schema registry, "Schema from ADR-019 event contract; compatibility mode BACKWARD"], status: done }
- { cells: [Switch the outbox relay from SQS to Kafka, "Behind flag orders_publish_kafka; dual publish for two weeks"], status: in progress }
- { cells: [Migrate Billing, Fulfilment, and Notifications consumers, "Billing done; the other two due 3 Oct"], status: in progress }
- { cells: [Consumer lag dashboard and page, "Page at lag over 5 minutes per group"], status: not started }
- { cells: [Delete the three SQS queues, "30 days after the last consumer moves"], status: not started }
```