Stage 1 of 9 · Model
Estimate the load
Before drawing any boxes, work out how much traffic and data this system really has. Rough numbers are enough: you only need to know whether something fits on one machine or needs many.
What you need to know first
Requirements usually come as monthly or daily totals. Systems fail per second, so convert.
A month has about 2.6 million seconds (30 days × 86,400 seconds). Divide a monthly total by 2.6 million to get the average per second. Traffic isn't flat, so multiply by the peak factor to get the rate you actually have to handle.
100 million new links a month. About how many links are created per second, on average?
About 38 per second.
100,000,000 ÷ 2,600,000 ≈ 38 per second. At 5× peak, about 190 per second.
One Postgres server handles thousands of small inserts a second, so writes are nowhere near a limit.
10 billion redirects a month, with peaks at 5× the average. About how many redirects per second at peak?
About 19,000 per second.
10,000,000,000 ÷ 2,600,000 ≈ 3,850 per second on average. × 5 ≈ 19,000 per second at peak.
That's about 100 reads for every write. This is a read-heavy system, so the redirect path is where the design effort goes.
Now storage. Estimate the size of one record, then multiply by how many you add per year.
A link row holds a 7-character code, the destination URL (usually 100 to 200 bytes, sometimes much longer), an owner id, timestamps and a status. With index overhead, 500 bytes is a reasonable round number. 100 million links a month is 1.2 billion a year.
1.2 billion links a year at 500 bytes each. About how much storage per year?
About 600 GB.
1,200,000,000 × 500 bytes = 600,000,000,000 bytes ≈ 600 GB a year.
A single database server can hold several terabytes, so this fits on one machine for years.
Compare your numbers with what one machine can do. These are rough figures for a well-provisioned server, worth remembering as orders of magnitude:
| Work | One server handles roughly |
|---|---|
| Postgres lookups by primary key | tens of thousands per second |
| Postgres small inserts | thousands per second |
| Redis gets | around 100,000 per second |
| Disk | several terabytes |
Put your three numbers (190 writes/s at peak, 19,000 reads/s at peak, 600 GB/year) next to that table. What do they tell you?
One database server can handle it, especially with a cache in front for popular links.
Every number is within one server's range. That means scale isn't the hard part of this problem. The hard parts are elsewhere: correct codes, distant users, counting clicks and takedowns.
What the stage asks
Where should the links be stored?
- Sound
One Postgres primary with a standby replica for failover, with the code as the primary key
It covers the numbers with room to spare, and the primary key gives you uniqueness for free. The standby protects against losing the server. You can add read replicas or partitioning later, when a measurement says you need them.
- Defensible
A Cassandra cluster partitioned by code
It would handle far more than this, but you pay to run a cluster, and you lose simple unique constraints (Cassandra's lightweight transactions can emulate them, at a latency cost). Pick it when writes or storage outgrow one machine. These numbers are about 100× away from that.
- Flawed
Redis as the only store, since every read is a key lookup
Lookups would be fast, but Redis keeps everything in memory (600 GB a year of RAM is expensive) and its usual persistence settings can lose the last second of writes in a crash. A lost link breaks every poster it's printed on. Redis works well as a cache in front of a durable store, not instead of one.
- Defensible
Postgres sharded across eight servers from day one
It works, and code-based sharding is a reasonable plan for later. Today it means eight servers to run and back up, and routing logic in the app, to hold data that fits on one.
What a strong answer covers
- Uses the rates: about 40 writes/s (190 at peak) and about 19,000 reads/s at peak.
- Uses the storage estimate: about 600 GB a year.
- Names what would justify distributing, such as storage beyond one machine or write rates beyond one primary.Supporting
The reasoning
- Divide monthly totals by 2.6 million to get per-second rates, then multiply by the peak factor.
- Compare each number with what one machine can do before you reach for a distributed database.
- When the numbers fit on one machine, the hard parts of the problem are correctness and latency, not scale.
Writes are about 40 a second, reads about 19,000 at peak, and storage grows about 600 GB a year. One database with a cache in front handles all of that.
So the rest of this investigation isn't about scaling. It covers four problems the numbers don't solve: generating codes correctly, serving users far from the servers, counting clicks without slowing redirects, and removing a bad link from every cache.