Skip to content

Design a Distributed Job Queue, stage 3 of 9: decide

A buffer that can hold a bad day

Thousands of job handlers and the worker fleet consume from Redis. Rewriting all of them at once would be a risky project on the most critical async path in the company.

System so far· 4 parts
12SERVICEWeb serversQUEUERedis queuesWORKERWorkersDATABASEDatabasesand services

Select a component to see what it is responsible for and which state it owns.

  1. 1Workers → Redis queues: Lease jobs
  2. 2Workers → Databases and services: Do the work

What you need to know

0 of 1 checks done
  1. A log like Kafka stores messages in order, on disk, and keeps them for a set time (say two days) whether or not anyone has read them. Consumers track how far they have read.

    So writing to the log never depends on consumers keeping up. Reading can lag by hours and the producer doesn't notice.

  2. Kafka and a Redis list hand out work differently:

    Kafka partitionRedis-style queue
    Unit of progressan offset: "everything up to here is done"each job, leased and acknowledged on its own
    A slow jobholds up the jobs behind it in its partitionholds up only its own worker
    Retrying one jobneeds extra machinery (retry topics)built in: let the lease expire or re-enqueue
  3. Check

    A Kafka consumer reads jobs 1 to 10 from a partition. Job 3 fails and must be retried later; 4 to 10 succeed. What offset can it commit?