System design · Queues, provider rate limits, retries
How to design a notification system
A notification system sends the right message to the right user on the right channel - a push to a phone, an email, a text. Sending one notification is trivial. The hard part is that you don't control the last hop: Apple, Google, your email and SMS providers all cap how fast you can send, and the moment breaking news hits, you want to send twenty times faster than usual.
Updated · 7 min read
Requirements
Pin down the channels and the triggers first. Interviewers want to hear which notifications are urgent and which can wait.
- Functional: send push (through APNs and FCM), email and SMS; trigger notifications from events ("someone replied") and on a schedule (a daily digest).
- Functional: respect user preferences - opt-outs per channel and per category, and quiet hours in the user's time zone; render messages from templates.
- Functional, optional: delivery tracking (sent, delivered, opened) and analytics per campaign.
- Non-functional: accepting an event must be fast - well under a second - and every notification should arrive within about a minute, even during a spike.
- Non-functional: never send the same notification twice, and never lose one because a provider was busy.
Capacity estimates
These are assumptions, stated out loud. The point is the order of magnitude, and one number dominates: your provider's rate limit.
| Quantity | Assumption | Result |
|---|---|---|
| Event notifications | about 43M a day | ≈ 500/s average |
| Breaking news | a 20× spike for a few minutes | ≈ 10,000/s at peak |
| Scheduled digests | about 8.6M a day, spread evenly | ≈ 100/s, all day |
| Provider limit | your provider's rate limit, 2,000/s in this example | a hard ceiling on sends |
| Delivery log | ≈ 300 bytes per notification, kept 90 days | ≈ 15 GB a day ≈ 1.4 TB |
Three conclusions. Normal traffic of about 600 a second is well under the provider's limit, so the steady state is easy. The spike is five times over the limit, so no amount of your own hardware lets you send it all immediately. And the delivery log is modest - it fits a single key-value table.
API design
Other services call one internal endpoint. It accepts the request and returns immediately; delivery happens later.
POST /api/notifications
{ "idempotency_key": "reply-88231-u_7", "user_id": "u_7",
"category": "replies", "template": "new_reply",
"data": { "author": "Sam", "thread_id": "t_42" } }
→ 202 Accepted { "notification_id": "n_5f1c", "status": "queued" }
GET /api/notifications/{id}
→ 200 OK { "status": "delivered", "channel": "push", "attempts": 1 }
PUT /api/users/{id}/preferences
{ "push": true, "email": false, "quiet_hours": { "start": "22:00", "end": "07:00", "tz": "Asia/Kolkata" } }The response is 202, not 200: the notification is queued, not sent. The caller's idempotency_key lets it retry safely - a second request with the same key returns the first notification instead of creating a new one.
Data model
Every access is by a known key - a user, a notification, a template - so a key-value store such as DynamoDB fits well.
notifications
notification_id string partition key
idempotency_key string unique, for dedup
user_id string
channel string push | email | sms
status string queued | sent | delivered | failed
attempts int
created_at timestamp
preferences
user_id string partition key
channels map per channel and category opt-outs
quiet_hours map start, end, time zone
devices
user_id string partition key
device_token string sort key, APNs or FCM token
templates
template_id string partition key
locale string sort key
subject, body textRemove a device token as soon as the provider reports it invalid. Stale tokens waste your rate limit on sends that can never arrive.
High-level design
Separate accepting a notification from sending it. The two run at different speeds, and only one of them is under your control.
- The notification API validates the request, checks the idempotency key, and puts a message on a queue such as SQS - then returns 202. The caller waits only for this.
- A scheduler such as EventBridge Scheduler fires the daily digests and puts them on the same queue. Scheduled and event-driven notifications share one path.
- Workers pull from the queue. Each one checks preferences and quiet hours, renders the template, picks the user's devices, and calls the provider: APNs or FCM for push, SES for email, an SMS provider for texts.
- Workers record the outcome in the delivery log. Provider callbacks (delivered, bounced, invalid token) update it later.
Where it breaks
The naive design calls the push provider from the request path: the API receives an event and sends the push before replying. Each provider call takes a few hundred milliseconds, so every API thread waits on a server you don't own. When breaking news multiplies events 20×, two things fail at once: the API runs out of threads and starts timing out, and it sends 10,000 pushes a second to a provider that accepts 2,000. The provider rejects the rest, and those notifications are simply lost.
Adding API servers makes it worse. More servers send more requests a second, which only get rejected faster.
The fix is a queue between accepting and sending. The API acknowledges events onto the queue in milliseconds, whatever the volume. Workers drain the queue at a fixed rate sized just under the provider's limit - say 1,600 a second against a limit of 2,000 - so nothing is rejected. During the spike the queue grows instead of erroring. A five-second burst at 10,000 a second leaves about 42,000 messages waiting; with 1,600 a second going out and 600 still arriving, the backlog clears in about 40 seconds. The spike becomes a short delivery delay, not errors.
Sizing workers against the provider's limit
- Your provider's rate limit is the ceiling, not your worker count. Measure one worker's send rate, then run as many as fit under the limit with some headroom - 70 to 80% of it.
- Headroom covers retries and other services sharing the same account. Running right at the limit means every retry tips you over.
- Watch queue age, not just queue depth. Alarm when the oldest message is older than your delivery target; that's the signal workers can't keep up.
Retries, backoff and dedup
- When a provider throttles or fails, retry with exponential backoff and jitter, so retries spread out instead of arriving as a second spike.
- SQS delivers at least once, so a worker can see the same message twice. Before sending, record the notification ID with a conditional write; if it's already there, skip it.
- After a fixed number of attempts, move the message to a dead-letter queue for inspection instead of retrying forever.
- Don't retry permanent errors - an invalid token or a hard email bounce. Mark them failed and clean up the token or address.
Preferences, quiet hours and priority
Check preferences in the worker, at send time, not when the event is created - a user may have opted out in between. If a notification lands in quiet hours, hold it: schedule it for the end of the window with EventBridge Scheduler or an SQS delay, rather than dropping it. Give urgent notifications (security alerts, one-time codes) their own queue so a marketing backlog never delays them.
What interviewers look for
- Keeping provider calls off the request path, and acknowledging after a durable queue.
- Treating the provider's rate limit as the real capacity of the system, and sizing workers just under it.
- Explaining how the queue turns a spike into delay, with numbers for how long the backlog takes to clear.
- Idempotency at both ends: an idempotency key from callers and dedup in workers, plus backoff and a dead-letter queue.
- Preferences, quiet hours and priority handled deliberately, not as an afterthought.
Frequently asked questions
Why not send notifications directly from the API?
+
Provider calls are slow and rate-limited. Calling them from the request path ties up API threads and, during a spike, sends more than the provider accepts, so notifications get rejected. A queue lets the API respond in milliseconds and lets workers send at a safe, steady rate.
How many workers should a notification system run?
+
Enough to stay just under your provider's rate limit, with headroom for retries. Measure how many sends one worker makes a second and divide the limit by that. Workers beyond the limit don't deliver more - they just get rejected faster.
How do you avoid sending the same notification twice?
+
Use idempotency at two points. Callers send an idempotency key so retried requests don't create new notifications. Workers record each notification ID with a conditional write before sending, because queues like SQS can deliver a message more than once.
How are scheduled notifications like daily digests handled?
+
A scheduler such as EventBridge Scheduler triggers them and puts them on the same queue as event-driven notifications. Workers treat both the same way, so preferences, rate limits and retries apply to every notification.
How do you track whether a notification was delivered?
+
Workers write a delivery record when they hand a notification to the provider. Providers report back later - SES through SNS notifications for bounces and deliveries, push providers through responses that flag invalid tokens - and those reports update the record.