Skip to content
chiltepin

Generated from: “Explain the notifications service to a new hire.”

Notifications service overview

AI-generated example · 2026-09-14. The source passes chiltepin check. System details and measurements are illustrative; review them before adapting this document.

DOCUMENTOnboarding

Notifications service — new hire guide

What the service does, who it talks to, and the words the team uses for it.

The notifications service turns events from other services into emails, SMS messages, and push notifications. It sends about 900,000 messages a day, and 70% of them are transactional: order confirmations, password resets, invoices. The service owns templates, user preferences, and delivery, and it owns nothing about why a message is sent. If Orders wants a new email, Orders publishes a new event, and this team adds a template.

The team is four engineers and one on-call rotation. You will ship your first template change to staging by the end of week one.

Who talks to whom

SECTION 01 · C4 model
C4 · CONTEXT
C4 diagram: 9 elements, 8 relationshipsSYSTEMOrdersPublishes order.placed andorder.shipped.SYSTEMBillingPublishes invoice.issued andpayment.failed.SYSTEMAuthCalls the send API directly forpassword resets.SYSTEMNotificationsRenders templates, checkspreferences, delivers, recordsthe outcome.DBPreferencesPostgres. Channel opt-ins peruser, quiet hours, locale.EXTSendGridEmail delivery and bouncewebhooks.EXTTwilioSMS delivery and statuscallbacks.EXTAPNs / FCMPush delivery to iOS andAndroid.PERSONCustomerReceives the message; editspreferences in the app.12345678
1publishes order events [Kafka]2publishes billing events [Kafka]3POST /v1/send [HTTPS]4reads opt-ins and locale5sends email; receives bounces [REST + webhook]6sends SMS; receives status [REST + webhook]7sends push [HTTP/2]8changes preferences [via the app]
LegendSYSTEMsoftware systemDBdatabaseEXTexternal systemPERSONpersonoutside the boundarydata storeuses

Two paths lead in. Kafka carries anything a service already emits as an event. The send API serves the few cases where a caller needs a synchronous answer, like a reset link. Everything else is fire-and-forget from the caller's point of view. The provider webhooks are the only way the service learns that an email bounced.

One message from event to inbox

SECTION 02 · Sequence
SEQUENCE
Sequence diagram: 12 messages between 5 actorsOrdersorder-eventsnotif-workerPreferencesEXTSendGridALT[202 Accepted][5xx or timeout]1order.placed (order_id, customer_id)2consume, group notifications3opt-ins and locale for customer_id4email: on, sms: off, locale: de-DE5render template order_placed in de-DE6POST /v3/mail/send7202 + message id8commit offset; write delivery row SENT950310retry with backoff, 5 attempts over 30 min11after 5 failures: park in notif-dlq12webhook: delivered or bounced (minutes later)
LegendcallresponseEXTexternal actorfragment (alt / opt / loop)active
Dedupe key: event_id + channelRetry budget: 5 attempts, 30 minutes, then the DLQ

A retry re-renders the template, so a template fix during an outage applies to the retried messages. The worker writes the delivery row as PENDING before the provider call. A crash between the call and the commit then produces a duplicate send, not a lost one. Duplicates are rare and the dedupe key stops most of them.

Who to ask

SECTION 03 · Team
DW
Dana Whitfield
Tech lead
Template engine, on-call schedule, this guide
RM
Ravi Menon
Backend
Kafka consumers, retries, the DLQ
SA
Sofia Alves
Backend
Provider integrations and webhooks
N
#notifications
Slack channel
Every question; median answer 15 minutes

Vocabulary

SECTION 04 · Glossary
Channel
one delivery medium: email, sms, or push. A user opts in per channel.
Template
a versioned Handlebars file per message type and locale, stored in the notif-templates repo.
Transactional
a message a user action caused; never blocked by quiet hours.
Campaign
a message marketing scheduled; respects quiet hours and a daily cap of 2 per user.
Delivery row
one Postgres row per message per channel with status PENDING, SENT, DELIVERED, BOUNCED, or FAILED.
DLQ
the notif-dlq topic; a message lands here after 5 failed provider calls and an on-call replays it.
Quiet hours
a per-user window, default 22:00 to 08:00 local, when campaigns wait.

Questions new engineers ask

SECTION 05 · FAQ
Why does the service not know why a message is sent?

So a product change never touches this service. Orders decides that a shipped order deserves an email; this team decides how the email looks and reaches the user.

What happens when SendGrid is down for an hour?

Retries cover 30 minutes, then messages park in the DLQ. The on-call replays the DLQ once the provider recovers. Nothing is lost; some messages are late.

Can I send a test message to myself?

Yes. Run make send-test TEMPLATE=order_placed TO=you@example.com against staging. Staging providers are sandboxes and deliver only to @example.com addresses.

Where do bounce and complaint rates live?

Grafana, dashboard notif-delivery. A bounce rate above 2% on any template pages the on-call, because it damages sender reputation for every template.

View the Markdown
```meta
title: Notifications service — new hire guide
subtitle: What the service does, who it talks to, and the words the team uses for it.
tag: Onboarding
```

The notifications service turns events from other services into emails, SMS messages, and push notifications. It sends about 900,000 messages a day, and 70% of them are transactional: order confirmations, password resets, invoices. The service owns templates, user preferences, and delivery, and it owns nothing about why a message is sent. If Orders wants a new email, Orders publishes a new event, and this team adds a template.

The team is four engineers and one on-call rotation. You will ship your first template change to staging by the end of week one.

## Who talks to whom

```c4
id: notif-context
level: context
nodes:
  - { id: orders, col: 1, row: 1, kind: system, name: Orders, desc: "Publishes order.placed and order.shipped." }
  - { id: billing, col: 1, row: 2, kind: system, name: Billing, desc: "Publishes invoice.issued and payment.failed." }
  - { id: auth, col: 1, row: 3, kind: system, name: Auth, desc: "Calls the send API directly for password resets." }
  - { id: notif, col: 2, row: 2, kind: system, name: Notifications, desc: "Renders templates, checks preferences, delivers, records the outcome." }
  - { id: prefs, col: 2, row: 3, kind: store, name: Preferences, desc: "Postgres. Channel opt-ins per user, quiet hours, locale." }
  - { id: sendgrid, col: 3, row: 1, kind: external, name: SendGrid, desc: "Email delivery and bounce webhooks." }
  - { id: twilio, col: 3, row: 2, kind: external, name: Twilio, desc: "SMS delivery and status callbacks." }
  - { id: apns, col: 3, row: 3, kind: external, name: "APNs / FCM", desc: "Push delivery to iOS and Android." }
  - { id: user, col: 4, row: 2, kind: person, name: Customer, desc: "Receives the message; edits preferences in the app." }
edges:
  - { from: orders, to: notif, label: "publishes order events", tech: Kafka }
  - { from: billing, to: notif, label: "publishes billing events", tech: Kafka }
  - { from: auth, to: notif, label: "POST /v1/send", tech: HTTPS }
  - { from: notif, to: prefs, label: "reads opt-ins and locale" }
  - { from: notif, to: sendgrid, label: "sends email; receives bounces", tech: REST + webhook }
  - { from: notif, to: twilio, label: "sends SMS; receives status", tech: REST + webhook }
  - { from: notif, to: apns, label: "sends push", tech: HTTP/2 }
  - { from: user, to: prefs, label: "changes preferences", tech: via the app }
```

Two paths lead in. Kafka carries anything a service already emits as an event. The send API serves the few cases where a caller needs a synchronous answer, like a reset link. Everything else is fire-and-forget from the caller's point of view. The provider webhooks are the only way the service learns that an email bounced.

## One message from event to inbox

```sequence
id: notif-one-message
actors:
  - { id: Orders, name: Orders }
  - { id: Kafka, name: order-events }
  - { id: Worker, name: notif-worker }
  - { id: Prefs, name: Preferences }
  - { id: SendGrid, name: SendGrid, external: true }
messages:
  - Orders -> Kafka: "order.placed (order_id, customer_id)"
  - Kafka --> +Worker: "consume, group notifications"
  - Worker -> Prefs: "opt-ins and locale for customer_id"
  - Prefs --> Worker: "email: on, sms: off, locale: de-DE"
  - Worker -> Worker: "render template order_placed in de-DE"
  - Worker -> +SendGrid: "POST /v3/mail/send"
  - alt: 202 Accepted
  - SendGrid --> -Worker: "202 + message id"
  - Worker -> Kafka: "commit offset; write delivery row SENT"
  - else: 5xx or timeout
  - SendGrid --> Worker: "503"
  - Worker -> Worker: "retry with backoff, 5 attempts over 30 min"
  - Worker -> Kafka: "after 5 failures: park in notif-dlq"
  - end
  - SendGrid --> -Worker: "webhook: delivered or bounced (minutes later)"
foot:
  - { label: Dedupe key, value: "event_id + channel" }
  - { label: Retry budget, value: "5 attempts, 30 minutes, then the DLQ" }
```

A retry re-renders the template, so a template fix during an outage applies to the retried messages. The worker writes the delivery row as PENDING before the provider call. A crash between the call and the commit then produces a duplicate send, not a lost one. Duplicates are rare and the dedupe key stops most of them.

## Who to ask

```team
id: notif-team
members:
  - { name: Dana Whitfield, role: Tech lead, focus: "Template engine, on-call schedule, this guide", accent: navy }
  - { name: Ravi Menon, role: Backend, focus: "Kafka consumers, retries, the DLQ", accent: teal }
  - { name: Sofia Alves, role: Backend, focus: "Provider integrations and webhooks", accent: purple }
  - { name: "#notifications", initials: N, role: Slack channel, focus: "Every question; median answer 15 minutes", accent: green }
```

## Vocabulary

```glossary
id: notif-terms
terms:
  - Channel — one delivery medium: email, sms, or push. A user opts in per channel.
  - Template — a versioned Handlebars file per message type and locale, stored in the notif-templates repo.
  - Transactional — a message a user action caused; never blocked by quiet hours.
  - Campaign — a message marketing scheduled; respects quiet hours and a daily cap of 2 per user.
  - Delivery row — one Postgres row per message per channel with status PENDING, SENT, DELIVERED, BOUNCED, or FAILED.
  - DLQ — the notif-dlq topic; a message lands here after 5 failed provider calls and an on-call replays it.
  - Quiet hours — a per-user window, default 22:00 to 08:00 local, when campaigns wait.
```

## Questions new engineers ask

```faq
id: notif-faq
items:
  - q: Why does the service not know why a message is sent?
    a: "So a product change never touches this service. Orders decides that a shipped order deserves an email; this team decides how the email looks and reaches the user."
  - q: What happens when SendGrid is down for an hour?
    a: "Retries cover 30 minutes, then messages park in the DLQ. The on-call replays the DLQ once the provider recovers. Nothing is lost; some messages are late."
  - q: Can I send a test message to myself?
    a: "Yes. Run make send-test TEMPLATE=order_placed TO=you@example.com against staging. Staging providers are sandboxes and deliver only to @example.com addresses."
  - q: Where do bounce and complaint rates live?
    a: "Grafana, dashboard notif-delivery. A bounce rate above 2% on any template pages the on-call, because it damages sender reputation for every template."
```