Skip to content

Slack: job queues under backlog

A job queue that failed because its backlog lived in memory, and the redesign that moved the backlog to disk.

The idea

Slack's job queue held every pending job in Redis. That was fine until a slow database made the workers slow. Jobs piled up, Redis filled, and because taking a job off the queue also needed a little free memory, the queue could no longer drain at all.

The redesign separates two things a queue does: holding a backlog, and handing out work. A durable log on disk took the backlog; Redis kept only the jobs workers were about to run; a relay between them set the pace. The rollout is as instructive as the design: run the new path next to the old one on real traffic, throw its output away, and switch only once it has proven itself.

Read the originals

Written by the engineers who built it.

  • Scaling Slack's Job Queue

    Saroj Yadav and others · Post, Dec 2017

    An outage, the design flaw behind it, and the redesign. The section on rolling the new path out without an outage is worth reading twice.

Practise it

Make the decisions yourself, then compare.