Design a Video Processing Pipeline, stage 4 of 14: decide
How do jobs reach workers?
A video is queued. Some process with CPU to spare has to find it, process it, and report back. Jobs take 5-20 minutes. You expect about 2,000 a day at launch and 10x within a year, which is still under one job every four seconds on average.
You have Postgres. A managed message queue (at-least-once delivery, visibility timeouts) is also available if you want one.
System so far· 5 parts
Select a component to see what it is responsible for and which state it owns.
- 1Instructor browser → Video API: Create upload, report parts, poll status
- 2Instructor browser → Object storage: Upload parts via presigned URLs
- 3Video API → Object storage: Complete multipart upload, verify object
- 4Reconciler → Postgres: Find abandoned uploads and orphaned output
- Request / response
- Bulk data
What you need to know
Work that outlives a request has to become a durable record somewhere: a row, or a message in a queue. A job that exists only in a process's memory dies with that process, and nothing remembers it was owed.
When one action changes two systems, say "mark the video queued" in Postgres and "publish a job" to a message queue, a crash between the two leaves them disagreeing. That's the dual write problem.
If both changes live in the same database, a single transaction makes them happen together or not at all.
Think first
The API updates the video to 'queued' in Postgres, then publishes a message to the queue. The process crashes between the two. What state is the system in?Postgres can act as a queue at modest volume. Workers claim jobs with:
SELECT id FROM jobs WHERE status = 'queued' ORDER BY created_at FOR UPDATE SKIP LOCKED LIMIT 1FOR UPDATElocks the chosen row;SKIP LOCKEDmakes other workers skip rows that are locked instead of waiting for them.Work it out
20,000 jobs a day (10× launch volume). About how many jobs per second is that on average?