Skip to content
chiltepin

Generated from: “Decide between Kafka and SQS for order events; write the decision record.”

Kafka vs SQS for order events decision record

AI-generated example · 2026-09-14. The source passes chiltepin check. System details and measurements are illustrative; review them before adapting this document.

DOCUMENTADR · Accepted

ADR-021 — Kafka vs SQS for order events

Which broker carries order.placed, and what the choice costs in ops and replay.

Orders publishes order.placed to three consumers today: Billing, Fulfilment, and Notifications. Fraud and Analytics join in Q4, and Fraud needs to replay 90 days of events to train its model. Volume is 40k orders a day, with a peak of 12 per second on campaign days. The current path is one SQS queue per consumer, and Orders writes each event three times. A fourth consumer means a fourth write in Orders code, and a missed write is silent.

Assumptions in this record: the platform runs on AWS. The team has no Kafka operator today. Every consumer already dedupes on event_id, so at-least-once delivery is acceptable on either broker.

Options

SECTION 01 · Options
Option 1SNS fan-out to SQS
Orders publishes once to an SNS topic; one SQS queue per consumer subscribes. FIFO topics for ordering.
  • No new infrastructure to run
  • Consumers keep their SQS client code
  • FIFO gives per-order ordering by message group
  • 14-day retention cap, so a 90-day replay needs a second copy in S3
  • FIFO throughput is 300 msg/s per group without batching
  • No consumer offsets: a replay means a re-publish from the archive
VIABLE — fallback if MSK Serverless pricing changes
Option 2Kafka on MSK Serverless
One topic, order-events, keyed by order_id. Each consumer runs its own consumer group.
  • Replay is a consumer group reset, no archive needed
  • Retention set to 90 days by config
  • Per-key ordering at any throughput
  • Schema registry enforces the event contract
  • New client library in five services
  • Consumer lag is a new alert to own
  • About $210 a month at our volume, vs $40 for SNS + SQS
CHOSEN
Option 3Self-managed Kafka on EC2
Three brokers across three zones, run by the platform team.
  • Lowest per-message cost at scale
  • Full control of broker config
  • Nobody on the team has run Kafka in production
  • Broker upgrades and rebalances become our pages
  • Break-even with MSK is above 2M events a day, 50× our volume
REJECTED — the ops cost exceeds the platform team's budget

Decision

SECTION 02 · Note
Note
Order events move to one Kafka topic, order-events, on MSK Serverless. Orders publishes once through the outbox relay. Each consumer owns a consumer group and its own offsets.

What we accepted

SECTION 03 · Trade-offs
We gain
Orders publishes once; a new consumer is a deploy in that team's repo
Fraud replays 90 days with a group reset in under an hour
Ordering per order_id holds at any throughput
The schema registry rejects a breaking payload change at publish time
We pay
Five services learn a new client and its rebalance behaviour
Consumer lag needs a dashboard and a page per group
$170 a month more than the SNS option
A consumer that falls 90 days behind loses events, where SQS would hold them for 14 days only

The trade we made explicit: $2,040 a year buys replay and a single publish path. One missed SQS write in the current design cost a two-day reconciliation in July. The team prefers a lag alert it can see to a missing event it cannot.

The chosen topology

SECTION 04 · Architecture
EVENT
Block diagram: 7 nodes, 6 connectionsConsumer groupsOrdersoutbox relayPRODUCERorder-eventsMSK Serverless · 90dTOPICBillingCONSUMERFulfilmentCONSUMERNotificationsCONSUMERFraudQ4CONSUMERAnalyticsQ4CONSUMER123456
1publish order.placed2charge the card3create the shipment4send the confirmation5score and replay6record the sale
LegendPRODUCERproducerTOPICtopicCONSUMERconsumercallsasync / optionalentry point

The topic is the only thing the two sides share. Orders holds no list of consumers and no per-consumer retry logic. A consumer that is down for a week reads every event it missed when it returns.

Follow-up work

SECTION 05 · Status
TaskUpdateStatus
Provision MSK Serverless and the order-events topicTerraform merged; 90-day retention setdone
Register order.placed v1 in the schema registrySchema from ADR-019 event contract; compatibility mode BACKWARDdone
Switch the outbox relay from SQS to KafkaBehind flag orders_publish_kafka; dual publish for two weeksin progress
Migrate BillingFulfilmentand Notifications consumersBilling done; the other two due 3 Octin progress
Consumer lag dashboard and pagePage at lag over 5 minutes per groupnot started
Delete the three SQS queues30 days after the last consumer movesnot started
Legenddonein progressnot started
View the Markdown
```meta
title: ADR-021 — Kafka vs SQS for order events
subtitle: Which broker carries order.placed, and what the choice costs in ops and replay.
tag: ADR · Accepted
```

Orders publishes `order.placed` to three consumers today: Billing, Fulfilment, and Notifications. Fraud and Analytics join in Q4, and Fraud needs to replay 90 days of events to train its model. Volume is 40k orders a day, with a peak of 12 per second on campaign days. The current path is one SQS queue per consumer, and Orders writes each event three times. A fourth consumer means a fourth write in Orders code, and a missed write is silent.

Assumptions in this record: the platform runs on AWS. The team has no Kafka operator today. Every consumer already dedupes on `event_id`, so at-least-once delivery is acceptable on either broker.

## Options

```options
id: adr021-options
items:
  - kicker: Option 1
    title: SNS fan-out to SQS
    how: Orders publishes once to an SNS topic; one SQS queue per consumer subscribes. FIFO topics for ordering.
    pros: [No new infrastructure to run, "Consumers keep their SQS client code", "FIFO gives per-order ordering by message group"]
    cons: ["14-day retention cap, so a 90-day replay needs a second copy in S3", "FIFO throughput is 300 msg/s per group without batching", "No consumer offsets: a replay means a re-publish from the archive"]
    verdict: "VIABLE — fallback if MSK Serverless pricing changes"
    tone: viable
  - kicker: Option 2
    title: Kafka on MSK Serverless
    how: One topic, order-events, keyed by order_id. Each consumer runs its own consumer group.
    pros: ["Replay is a consumer group reset, no archive needed", "Retention set to 90 days by config", "Per-key ordering at any throughput", "Schema registry enforces the event contract"]
    cons: ["New client library in five services", "Consumer lag is a new alert to own", "About $210 a month at our volume, vs $40 for SNS + SQS"]
    verdict: "CHOSEN"
    tone: chosen
  - kicker: Option 3
    title: Self-managed Kafka on EC2
    how: Three brokers across three zones, run by the platform team.
    pros: [Lowest per-message cost at scale, Full control of broker config]
    cons: ["Nobody on the team has run Kafka in production", "Broker upgrades and rebalances become our pages", "Break-even with MSK is above 2M events a day, 50× our volume"]
    verdict: "REJECTED — the ops cost exceeds the platform team's budget"
    tone: rejected
```

## Decision

```callout
tone: note
title: Decision
body: "Order events move to one Kafka topic, order-events, on MSK Serverless. Orders publishes once through the outbox relay. Each consumer owns a consumer group and its own offsets."
```

## What we accepted

```proscons
id: adr021-consequences
prosLabel: We gain
consLabel: We pay
pros:
  - Orders publishes once; a new consumer is a deploy in that team's repo
  - Fraud replays 90 days with a group reset in under an hour
  - Ordering per order_id holds at any throughput
  - The schema registry rejects a breaking payload change at publish time
cons:
  - Five services learn a new client and its rebalance behaviour
  - Consumer lag needs a dashboard and a page per group
  - "$170 a month more than the SNS option"
  - A consumer that falls 90 days behind loses events, where SQS would hold them for 14 days only
```

The trade we made explicit: $2,040 a year buys replay and a single publish path. One missed SQS write in the current design cost a two-day reconciliation in July. The team prefers a lag alert it can see to a missing event it cannot.

## The chosen topology

```block
id: adr021-topology
preset: event
groups:
  - { id: consumers, col: 3, row: 1, cols: 1, rows: 5, label: Consumer groups }
nodes:
  - { id: orders, col: 1, row: 3, kind: producer, name: Orders, tech: outbox relay }
  - { id: topic, col: 2, row: 3, kind: topic, name: order-events, tech: "MSK Serverless · 90 d" }
  - { id: billing, col: 3, row: 1, kind: consumer, name: Billing }
  - { id: fulfilment, col: 3, row: 2, kind: consumer, name: Fulfilment }
  - { id: notifications, col: 3, row: 3, kind: consumer, name: Notifications }
  - { id: fraud, col: 3, row: 4, kind: consumer, name: Fraud, tech: Q4 }
  - { id: analytics, col: 3, row: 5, kind: consumer, name: Analytics, tech: Q4 }
edges:
  - orders -> topic: publish order.placed
  - topic --> billing: charge the card
  - topic --> fulfilment: create the shipment
  - topic --> notifications: send the confirmation
  - topic --> fraud: score and replay
  - topic --> analytics: record the sale
```

The topic is the only thing the two sides share. Orders holds no list of consumers and no per-consumer retry logic. A consumer that is down for a week reads every event it missed when it returns.

## Follow-up work

```statustable
id: adr021-followups
columns: [Task, Update]
statuses:
  - { label: done, color: success }
  - { label: in progress, color: blue }
  - { label: not started, color: neutral }
rows:
  - { cells: [Provision MSK Serverless and the order-events topic, "Terraform merged; 90-day retention set"], status: done }
  - { cells: [Register order.placed v1 in the schema registry, "Schema from ADR-019 event contract; compatibility mode BACKWARD"], status: done }
  - { cells: [Switch the outbox relay from SQS to Kafka, "Behind flag orders_publish_kafka; dual publish for two weeks"], status: in progress }
  - { cells: [Migrate Billing, Fulfilment, and Notifications consumers, "Billing done; the other two due 3 Oct"], status: in progress }
  - { cells: [Consumer lag dashboard and page, "Page at lag over 5 minutes per group"], status: not started }
  - { cells: [Delete the three SQS queues, "30 days after the last consumer moves"], status: not started }
```