A notification system that survives a traffic spike
Flash sale, millions of orders, SMS and email on the critical path. Why the Order API should stay small, how an outbox and a queue absorb the burst, and what still breaks if workers cannot catch up.
A notification system that survives a traffic spike
Millions of orders in minutes.
Your notification system can't process them all instantly.
So what happens to the rest?
That's the part people skip. Not “how do I send an SMS.” What do you do with the work that doesn't fit in this second?
Flash sale at 10:00. Orders pile in. Each one might mean push, email, SMS, WhatsApp, a payment update. If the Order API waits for all of that before it returns, you've tied “we stored the order” to “Twilio is having a good day.”
I've seen that look fine in staging. Then checkout looks down because a provider sat at eight seconds. The order was already valid. Nobody needed to wait for the receipt email to exist.
The first drawing that hurts
Order API → SMS → Email → Push → Response
One slow hop holds the request. Timeouts stack. The pool empties. Healthy orders wait behind dead SMS calls. You didn't have an order problem. You had a notification problem wearing the Order API's clothes.
More API pods won't fix that. The API should do less.
┌───────────────┐
Client ────────→ │ Order API │
└───────┬───────┘
│ same DB transaction
▼
┌───────────────┐
│ DB │
│ order + │
│ outbox │
└───────┬───────┘
│
Publisher
│
▼
┌───────────────┐
│ Kafka │
└───────┬───────┘
│
┌──────────────┼──────────────┐
▼ ▼ ▼
Push worker Email worker SMS worker
│ │ │
▼ ▼ ▼
Provider Provider Provider
Critical path: write the order. Everything else leaves the request.
Keep the critical path small
The Order API validates, writes the order, and writes a notification event in the same database transaction — an outbox row. Then it returns.
A publisher (poller or CDC) reads outbox rows and publishes them to Kafka. After a successful publish, it records that the row was published. That last step can fail. The publisher can crash after Kafka accepted the message and before the outbox is updated. Then the same event goes out again. That's expected. Downstream has to live with duplicates.
Why the extra table? If you COMMIT the order and then produce() to Kafka as a second step, a crash between those two loses the event. If you publish first and the database transaction later rolls back, consumers may process an event for an order that was never committed. Two independent writes is how things go missing.
The outbox is boring. It gives you a durable record of the *intent* to publish, instead of hoping both writes succeed together. It does not, by itself, mean the SMS landed. That's a later problem.
I wouldn't put Kafka inside the request either. Then Kafka latency becomes checkout latency. The request already did its job: the order is durable, and so is the intent to notify.
Let a queue absorb the spike — with a catch
Say workers and providers can finish about 500K notifications per minute, and 5M show up in a short burst.
That extra 4.5M does not belong in API memory, and it does not belong as 4.5 million threads hitting Twilio. A log like Kafka (or a job queue, at smaller scale) holds the backlog. Consumers pull what they can. The Order API stays up.
Kafka does not delete the spike. It stores it. Storage is not infinite either — retention has a limit, disks fill. If producers keep outrunning consumers, lag grows. The sale still “works” while receipts are hours late. Fine for some email. Not fine if you tell yourself the queue made the problem go away.
During the spike, “is my API fast?” is the easy dashboard. The useful one is: is the backlog growing faster than workers can drain it?
If inflow stays above drain rate, you need more workers, slower publish, dropping a low-priority channel, or a product call that SMS can wait. You don't need a second Kafka cluster because it looks serious.
A lot of things I've shipped were fine on Redis plus a job library (BullMQ and friends). Kafka starts to earn the ops cost when several teams consume the same events, you want a durable log, and the volume is actually there.
Isolate channels
Email, SMS, push, and WhatsApp fail for different reasons, at different rates. One worker that does all four in a loop means an SMS outage parks push too.
Separate consumers — or at least separate queues — per channel. SMS is down: push still moves. You can add SMS workers without buying email capacity you don't need.
SLAs differ too. Push can be “soon.” A receipt email can be minutes. OTP SMS is another product. I wouldn't dump OTPs onto the same pile as marketing WhatsApp and hope.
The cost is real: more things to deploy, more lag graphs, more pages at night. Isolation is for failure domains, not for a résumé.
Retry carefully, then give up into a DLQ
Retries help a 503. They also send a second SMS if you aren't [idempotent](/insights/idempotency-keys-payments).
Backoff with jitter, so a recovering provider isn't hit by every worker at once. A max attempt count, then a dead-letter queue. Retry forever and you'll DDoS your own vendor while the poison message sits there unread.
A DLQ with no alert is how “we have retries” turns into “we lost 40K payment SMS overnight.”
Idempotency — ours and theirs
A common Kafka consumer setup gives you at-least-once processing. Semantics depend on how you commit offsets, whether you use transactions, and what you do *outside* Kafka. In the setups most of us run: the worker can crash after the provider accepts the SMS and before the offset is committed. You'll see the message again.
Give every notification a stable id — order + channel + purpose is enough (order_123:sms:placed).
A unique constraint on our side stops us from treating the same event as two jobs. It does not make an external SMS exactly-once.
Mark “delivered” before you send, then crash: you may never send.
mark delivered
→ send SMS
→ worker dies
Send first, then crash before you record it: the retry can send twice.
send SMS
→ worker dies
→ “delivered” never written
That's the gap. Our database can dedupe *our* processing. It can't unsay a text that already left the provider. When the vendor supports idempotency keys, pass that same stable id so *they* can ignore the retry too.
Exactly-once across an SMS network is a nice slide. Don't design as if it's true.
Protect providers
Vendors rate-limit. Burst 5M SMS and you get 429s, then a worse outage than the one you had.
Cap calls per provider in the worker. Circuit breakers: after a stretch of failures, stop calling for a bit, fail into retry or the DLQ, probe later. Don't hammer a dead region.
That's not manners. That's how one bad SMS integration doesn't burn the connections push still needs.
Watch the bottleneck that matters
API p95 can look green while lag climbs.
I'd watch:
- consumer lag, per channel
- delivery success rate
- retry count
- DLQ size
- provider latency
Lag going up during the sale means you're borrowing time. Scale workers, or shed work (delay WhatsApp, keep payment SMS). Graphs without that decision are just graphs.
One order, in real time
10:00:01 — they place an order. API writes orders and outbox in one transaction. Response in tens of milliseconds. No Twilio.
10:00:02 — publisher publishes OrderPlaced. If it dies before it records that, Kafka may get the event twice. Workers have to handle that.
10:00:03 — push goes out. SMS workers are behind. Push arrived. SMS is still in the log, not in the API.
10:04 — SMS catches up. Same id: skip if we already recorded a send; if not, the provider key is what stops a double text.
10:30 — if SMS is still failing, those ids sit in a DLQ with a reason. Checkout never knew.
When I'd use this
Orders or payments where losing the write is worse than delaying the ping. Spiky traffic. More than one channel. Third parties you don't control.
Durable intent, async drain, isolation, idempotency on both sides of the wire, lag you can see.
When I wouldn't
200 emails a day. An admin “notify me” button. One channel, a provider that's usually fine.
Insert a job row, one worker, retries, done. Kafka, an outbox poller, four consumer groups, and a DLQ UI is a lot of machinery. I've seen that cost more than the spike it was supposed to survive — because the spike never showed up.
Don't dump OTP and receipt on the same pile either. Different urgency. Different failure budget.
Interviews vs the pager
In a design interview they want the split: sync money/order, async fan-out, plus lag, retries, duplicates. Drawing Kafka isn't the point. Saying what happens when consumers are slower than producers is.
In production you also care who owns the lag graph at 10:05, whether the outbox publisher is a single box, and whether the DLQ has a runbook. A design the team can't run will lose messages quietly.
What to remember
Keep the critical path small. Store the order. Notifications can wait — if the wait is bounded, visible, and something you can actually drain.
A queue can hold the rest. It can't process faster than your workers and providers. If lag only grows, you don't have a messaging problem. You have a capacity problem. The log is just honest about it.