Design a Collaborative Editor (Google Docs), stage 9 of 12: break it
Deploy day
Deploys are the most common "failure" this system will ever see, and they happen every day. Evaluate each statement about getting through one.
System so far· 4 parts
Select a component to see what it is responsible for and which state it owns.
- 1Editor client → Document router: WebSocket: ops, acks, remote ops, presence
- 2Document router → Document owner: Route by document id to the current owner
- 3Document owner → Postgres op log: Append ops at next seq (batched, epoch-fenced)
- Server push
- Request / response
What you need to know
Load-balancer draining waits for in-flight requests to finish before stopping a server. A WebSocket is one request that never finishes, so draining alone just waits out the timeout and cuts it.
The server has to hand off actively: stop accepting ops, flush pending batches, release ownership, and tell clients to reconnect.
Work it out
20,000 clients reconnect with a random delay spread evenly over 10 seconds. About how many reconnections a second do the new servers face?During handoff, two servers can briefly both believe they own a document: the old one mid-flush, the new one starting. Ownership is a lease with an epoch number that increases with each new owner.
Appends to the log are conditioned on the owner's epoch, and
(doc_id, seq)is unique, so a stale owner's write fails and it learns it lost.Check
Should presence be saved before a deploy so avatars survive it?