Notifications across email, push and in-app
Design a Notification System, from a blank page
This is how the interview actually runs: one prompt, and you decide what to cover and in what order. Write each section, then compare it with a reference design and see what you left out.
A 45-minute round. You drive; nothing prompts you.
The prompt
You own notifications for a project-management product with five million users. People are notified when they are mentioned, assigned an issue, or someone comments on something they follow. Security events (a new login, a password change) must reach them quickly. Marketing occasionally announces features to everyone.
Notifications go out by email (through a provider with a 500 sends/second account limit), mobile push (APNs and FCM) and an in-app inbox. Users choose, per notification type and channel, what they want, and set quiet hours.
Today, each product service sends its own emails inline. Last week an email-provider outage made commenting fail for twenty minutes, and users regularly complain about duplicate notifications.
What the interviewer would tell you if you asked
- About 2 million product events a day fan out to ~8 million notifications, with work-hour peaks around 10x the average rate.
- Email provider: 500 sends/second account-wide, no idempotency keys, responses take 50 ms to several seconds.
- Push providers accept high throughput but report invalid device tokens per send.
- A managed queue, Postgres and Redis are available; product services already write events through a transactional outbox.
- Each product event has a unique, stable id.
- Users may have several devices, each with its own push token.
- Legal unsubscribe handling for marketing email is required in addition to in-app preferences.
01
about 5 minWhat does the system have to do, and how well? List the functional requirements, then the non-functional ones (latency, availability, consistency, scale), and the questions you would ask the interviewer.
02
about 5 minTurn the volumes into the numbers that drive the design: requests per second at peak, storage, bandwidth, and anything else that decides whether one machine is enough.
03
about 10 minName the components and what each one is responsible for. Then trace the main request through them, and say where the durable state lives.
04
about 15 minPick the hardest decisions in this design and make them: what you chose, what you rejected, and which constraint decided it.
05
about 10 minWhat breaks? Walk through crashes, duplicates, slow dependencies and overload, and what the design does in each. Then: what changes at ten times the load, or with a new requirement?
Write something in at least 3 sections first. Gaps are fine; the comparison shows what they cost.