Skip to content

Design a Collaborative Editor (Google Docs), stage 7 of 12: break it

Acknowledged, then lost

To reduce database load, an engineer changed the owner to buffer ops in memory and flush them to Postgres every two seconds. Find the design decisions that turned a crash into data loss.

System so far· 4 parts
123CLIENTEditor clientEDGEDocument routerSERVICEDocument ownerDATABASEPostgres op log

Select a component to see what it is responsible for and which state it owns.

  1. 1Editor client → Document router: WebSocket: ops, acks, remote ops, presence
  2. 2Document router → Document owner: Route by document id to the current owner
  3. 3Document owner → Postgres op log: Append ops at next seq (batched, epoch-fenced)
  • Server push
  • Request / response

What you need to know

0 of 2 checks done
  1. Group commit batches many writes into one transaction without weakening any of them: ops for a document accumulate for 10–20 ms, one transaction inserts them all, and then each op is acknowledged.

    Throughput is the same as buffering longer; latency rises by a few milliseconds; and an ack still means "durable".

  2. Check

    The owner acks each op immediately and flushes to Postgres every 2 seconds. The process is killed 1.5 s after the last flush. What's lost?