Shard a Live Database Without Downtime, stage 1 of 9: model
What is actually running out?
Postgres tags each transaction with a 32-bit ID. Vacuum must periodically "freeze" old rows so IDs can be reused; if it falls too far behind, Postgres stops accepting writes rather than risk corrupting data.
System so far· 3 parts
Select a component to see what it is responsible for and which state it owns.
- 1Users → Application servers: Requests
- 2Application servers → Monolith Postgres: Reads and writes until cutover
What you need to know
Postgres gives each transaction a 32-bit transaction ID. That's about 4 billion IDs, reused in a cycle. To make reuse safe, vacuum must "freeze" old rows so they no longer depend on their original transaction ID.
If vacuum falls too far behind, Postgres stops accepting writes rather than risk misreading old rows as new. That's wraparound, and it's a hard deadline.
Three different fixes target three different limits:
Fix What it divides Read replicas read load (every replica still replays every write) Partitioning tables on one host maintenance work per table (same host limits) Sharding across hosts write volume, table size and vacuum work See Partitioning and Replication.
Check
Vacuum can't keep up with the primary's write volume. Do read replicas help?