Design a Distributed Job Queue, stage 3 of 9: decide
A buffer that can hold a bad day
Thousands of job handlers and the worker fleet consume from Redis. Rewriting all of them at once would be a risky project on the most critical async path in the company.
System so far· 4 parts
Select a component to see what it is responsible for and which state it owns.
- 1Workers → Redis queues: Lease jobs
- 2Workers → Databases and services: Do the work
What you need to know
A log like Kafka stores messages in order, on disk, and keeps them for a set time (say two days) whether or not anyone has read them. Consumers track how far they have read.
So writing to the log never depends on consumers keeping up. Reading can lag by hours and the producer doesn't notice.
Kafka and a Redis list hand out work differently:
Kafka partition Redis-style queue Unit of progress an offset: "everything up to here is done" each job, leased and acknowledged on its own A slow job holds up the jobs behind it in its partition holds up only its own worker Retrying one job needs extra machinery (retry topics) built in: let the lease expire or re-enqueue Check
A Kafka consumer reads jobs 1 to 10 from a partition. Job 3 fails and must be retried later; 4 to 10 succeed. What offset can it commit?When a change touches the most critical path in a system, how you get there matters as much as where you end up. A design that keeps thousands of existing workers unchanged can be rolled out one job type at a time and rolled back at any point. A design that rewrites them all has to work everywhere on the first try.