Design a Collaborative Editor (Google Docs), stage 7 of 12: break it
Acknowledged, then lost
To reduce database load, an engineer changed the owner to buffer ops in memory and flush them to Postgres every two seconds. Find the design decisions that turned a crash into data loss.
System so far· 4 parts
Select a component to see what it is responsible for and which state it owns.
- 1Editor client → Document router: WebSocket: ops, acks, remote ops, presence
- 2Document router → Document owner: Route by document id to the current owner
- 3Document owner → Postgres op log: Append ops at next seq (batched, epoch-fenced)
- Server push
- Request / response
What you need to know
0 of 2 checks done
Group commit batches many writes into one transaction without weakening any of them: ops for a document accumulate for 10–20 ms, one transaction inserts them all, and then each op is acknowledged.
Throughput is the same as buffering longer; latency rises by a few milliseconds; and an ack still means "durable".
Check
The owner acks each op immediately and flushes to Postgres every 2 seconds. The process is killed 1.5 s after the last flush. What's lost?Think first
After the crash, Bob reconnects claiming last_seq = 5131, but the log only reaches 5120. What should the new owner do?