Stage 1 of 9 · Model
What the numbers say
Use 86,400 seconds in a day. Messages are about 1 KB.
What you need to know first
120 million messages a day. About how many writes a second on average?
About 1,400 per second.
120,000,000 ÷ 86,400 ≈ 1,400 a second. Even several times that at peak is within one good database's reach. The write rate isn't the problem.
At about 1 KB per message, how many terabytes a year (before replication)?
About 44 TB.
120 GB a day × 365 ≈ 44 TB a year, and growing several-fold a year. What outgrows one machine is the data.
When data and indexes fit in memory, a random read is a memory lookup. When they don't, it becomes a disk seek, and latency becomes unpredictable. The usual fix is to store what one query needs physically together, so it takes one seek and a short sequential read instead of fifty random ones.
A Snowflake ID is 64 bits: a millisecond timestamp in the high bits, then a worker number and a per-worker sequence. Any server can generate one without coordinating, and sorting IDs sorts messages by creation time. See Generating unique identifiers.
So "the latest 50 messages in a channel" is "the 50 largest IDs in that channel".
Do time-ordered IDs require one central counter?
No: putting the timestamp in the high bits makes independently generated IDs sort by time.
Each worker's number and sequence keep IDs unique; the timestamp makes them sortable. No coordination needed.
What the stage asks
Which statements follow from the scenario?
- Holds
120 million messages a day is roughly 1,400 writes a second on average.
120,000,000 / 86,400 ≈ 1,390. Peaks are several times that, which a single good database could still absorb. The write rate is not the problem.
- Holds
What outgrows one machine is the data, not the request rate.
At ~1 KB a message that is ~120 GB a day, over 40 TB a year before replication, and growing. Once data and indexes stop fitting in memory, random reads go to disk and latency becomes unpredictable, which is exactly the symptom in the scenario.
- Depends
Because reads are about half the traffic, a cache of recent messages will take most read load off the database.
It helps for the latest page of busy channels. But reads are scattered across millions of small channels, history pages, jumps to old messages and mentions, so much of the read traffic is a long tail a cache will mostly miss. The store itself must serve random reads well.
- Holds
The dominant query is 'latest messages in channel X', so a channel's messages should be stored together, in time order.
If a page of messages is physically contiguous, loading it is one seek and a short sequential read rather than fifty random lookups.
- Fails
To sort messages by time, IDs must come from one central sequence.
Snowflake IDs put a millisecond timestamp in the high bits, then a worker number and a per-worker sequence. Any server generates them without coordination, and they sort by creation time to the millisecond. See Generating unique identifiers.
The reasoning
- The write rate is modest; what outgrows one machine is the ever-growing data.
- Store a channel's messages together in time order so the latest page is one contiguous read.
- Snowflake IDs sort by time without central coordination.
Chat storage is a data-shape problem. The write rate is modest; the dataset is huge, ever-growing and read at random. The design has to make the common read (a channel's latest page) touch a small, contiguous piece of data, whatever the total size.
Time-sortable IDs are what make that possible: if the ID is the time, "the latest 50" is just "the 50 largest IDs in this channel".