Design a Distributed Job Queue, stage 2 of 9: break it
Why the queue stopped draining
Select the lines that describe causes in the design, not just symptoms.
System so far· 4 parts
Select a component to see what it is responsible for and which state it owns.
- 1Workers → Redis queues: Lease jobs
- 2Workers → Databases and services: Do the work
What you need to know
Redis has a memory limit (
maxmemory). When a Redis used as a queue reaches it, commands that would add data fail with an out-of-memory error. Commands that only remove data still work.The catch is in the details: a reliable dequeue usually moves the job to a "processing" list (
RPOPLPUSH) so it isn't lost if the worker crashes. Moving writes a new entry, and writing needs memory.Think first
Redis is at its memory limit. Workers dequeue with RPOPLPUSH, which writes the job into a processing list. What happens to the queue?When a shared pool of workers serves several job types, each worker takes whatever job is next. Fast jobs leave quickly; slow jobs stay. Over time, the workers fill up with whichever type is slowest.
Work it out
200 workers share a queue. Search jobs used to take 30 ms; now they take 1.1 s. If 10% of arriving jobs are search jobs and the rest take 30 ms, roughly what share of busy worker time goes to search?Check
Every worker holds a connection to every Redis instance. During the outage, operators add 200 workers. What does that do?