System design · Queues, provider rate limits, retries

How to design a notification system

A notification system sends the right message to the right user on the right channel - a push to a phone, an email, a text. Sending one notification is trivial. The hard part is that you don't control the last hop: Apple, Google, your email and SMS providers all cap how fast you can send, and the moment breaking news hits, you want to send twenty times faster than usual.

Updated · 7 min read

Requirements

Pin down the channels and the triggers first. Interviewers want to hear which notifications are urgent and which can wait.

  • Functional: send push (through APNs and FCM), email and SMS; trigger notifications from events ("someone replied") and on a schedule (a daily digest).
  • Functional: respect user preferences - opt-outs per channel and per category, and quiet hours in the user's time zone; render messages from templates.
  • Functional, optional: delivery tracking (sent, delivered, opened) and analytics per campaign.
  • Non-functional: accepting an event must be fast - well under a second - and every notification should arrive within about a minute, even during a spike.
  • Non-functional: never send the same notification twice, and never lose one because a provider was busy.

Capacity estimates

These are assumptions, stated out loud. The point is the order of magnitude, and one number dominates: your provider's rate limit.

QuantityAssumptionResult
Event notificationsabout 43M a day≈ 500/s average
Breaking newsa 20× spike for a few minutes≈ 10,000/s at peak
Scheduled digestsabout 8.6M a day, spread evenly≈ 100/s, all day
Provider limityour provider's rate limit, 2,000/s in this examplea hard ceiling on sends
Delivery log≈ 300 bytes per notification, kept 90 days≈ 15 GB a day ≈ 1.4 TB

Three conclusions. Normal traffic of about 600 a second is well under the provider's limit, so the steady state is easy. The spike is five times over the limit, so no amount of your own hardware lets you send it all immediately. And the delivery log is modest - it fits a single key-value table.

API design

Other services call one internal endpoint. It accepts the request and returns immediately; delivery happens later.

POST /api/notifications
{ "idempotency_key": "reply-88231-u_7", "user_id": "u_7",
  "category": "replies", "template": "new_reply",
  "data": { "author": "Sam", "thread_id": "t_42" } }
→ 202 Accepted  { "notification_id": "n_5f1c", "status": "queued" }

GET /api/notifications/{id}
→ 200 OK  { "status": "delivered", "channel": "push", "attempts": 1 }

PUT /api/users/{id}/preferences
{ "push": true, "email": false, "quiet_hours": { "start": "22:00", "end": "07:00", "tz": "Asia/Kolkata" } }

The response is 202, not 200: the notification is queued, not sent. The caller's idempotency_key lets it retry safely - a second request with the same key returns the first notification instead of creating a new one.

Data model

Every access is by a known key - a user, a notification, a template - so a key-value store such as DynamoDB fits well.

notifications
  notification_id   string     partition key
  idempotency_key   string     unique, for dedup
  user_id           string
  channel           string     push | email | sms
  status            string     queued | sent | delivered | failed
  attempts          int
  created_at        timestamp

preferences
  user_id           string     partition key
  channels          map        per channel and category opt-outs
  quiet_hours       map        start, end, time zone

devices
  user_id           string     partition key
  device_token      string     sort key, APNs or FCM token

templates
  template_id       string     partition key
  locale            string     sort key
  subject, body     text

Remove a device token as soon as the provider reports it invalid. Stale tokens waste your rate limit on sends that can never arrive.

High-level design

Separate accepting a notification from sending it. The two run at different speeds, and only one of them is under your control.

  • The notification API validates the request, checks the idempotency key, and puts a message on a queue such as SQS - then returns 202. The caller waits only for this.
  • A scheduler such as EventBridge Scheduler fires the daily digests and puts them on the same queue. Scheduled and event-driven notifications share one path.
  • Workers pull from the queue. Each one checks preferences and quiet hours, renders the template, picks the user's devices, and calls the provider: APNs or FCM for push, SES for email, an SMS provider for texts.
  • Workers record the outcome in the delivery log. Provider callbacks (delivered, bounced, invalid token) update it later.

Where it breaks

The naive design calls the push provider from the request path: the API receives an event and sends the push before replying. Each provider call takes a few hundred milliseconds, so every API thread waits on a server you don't own. When breaking news multiplies events 20×, two things fail at once: the API runs out of threads and starts timing out, and it sends 10,000 pushes a second to a provider that accepts 2,000. The provider rejects the rest, and those notifications are simply lost.

Adding API servers makes it worse. More servers send more requests a second, which only get rejected faster.

The fix is a queue between accepting and sending. The API acknowledges events onto the queue in milliseconds, whatever the volume. Workers drain the queue at a fixed rate sized just under the provider's limit - say 1,600 a second against a limit of 2,000 - so nothing is rejected. During the spike the queue grows instead of erroring. A five-second burst at 10,000 a second leaves about 42,000 messages waiting; with 1,600 a second going out and 600 still arriving, the backlog clears in about 40 seconds. The spike becomes a short delivery delay, not errors.

Sizing workers against the provider's limit

  • Your provider's rate limit is the ceiling, not your worker count. Measure one worker's send rate, then run as many as fit under the limit with some headroom - 70 to 80% of it.
  • Headroom covers retries and other services sharing the same account. Running right at the limit means every retry tips you over.
  • Watch queue age, not just queue depth. Alarm when the oldest message is older than your delivery target; that's the signal workers can't keep up.

Retries, backoff and dedup

  • When a provider throttles or fails, retry with exponential backoff and jitter, so retries spread out instead of arriving as a second spike.
  • SQS delivers at least once, so a worker can see the same message twice. Before sending, record the notification ID with a conditional write; if it's already there, skip it.
  • After a fixed number of attempts, move the message to a dead-letter queue for inspection instead of retrying forever.
  • Don't retry permanent errors - an invalid token or a hard email bounce. Mark them failed and clean up the token or address.

Preferences, quiet hours and priority

Check preferences in the worker, at send time, not when the event is created - a user may have opted out in between. If a notification lands in quiet hours, hold it: schedule it for the end of the window with EventBridge Scheduler or an SQS delay, rather than dropping it. Give urgent notifications (security alerts, one-time codes) their own queue so a marketing backlog never delays them.

What interviewers look for

  • Keeping provider calls off the request path, and acknowledging after a durable queue.
  • Treating the provider's rate limit as the real capacity of the system, and sizing workers just under it.
  • Explaining how the queue turns a spike into delay, with numbers for how long the backlog takes to clear.
  • Idempotency at both ends: an idempotency key from callers and dedup in workers, plus backoff and a dead-letter queue.
  • Preferences, quiet hours and priority handled deliberately, not as an afterthought.

Frequently asked questions

Why not send notifications directly from the API?

+

Provider calls are slow and rate-limited. Calling them from the request path ties up API threads and, during a spike, sends more than the provider accepts, so notifications get rejected. A queue lets the API respond in milliseconds and lets workers send at a safe, steady rate.

How many workers should a notification system run?

+

Enough to stay just under your provider's rate limit, with headroom for retries. Measure how many sends one worker makes a second and divide the limit by that. Workers beyond the limit don't deliver more - they just get rejected faster.

How do you avoid sending the same notification twice?

+

Use idempotency at two points. Callers send an idempotency key so retried requests don't create new notifications. Workers record each notification ID with a conditional write before sending, because queues like SQS can deliver a message more than once.

How are scheduled notifications like daily digests handled?

+

A scheduler such as EventBridge Scheduler triggers them and puts them on the same queue as event-driven notifications. Workers treat both the same way, so preferences, rate limits and retries apply to every notification.

How do you track whether a notification was delivered?

+

Workers write a delivery record when they hand a notification to the provider. Providers report back later - SES through SNS notifications for bounces and deliveries, push providers through responses that flag invalid tokens - and those reports update the record.

Now break one yourself.

The first challenge takes about two minutes. No signup.