Skip to content

Design a Video Processing Pipeline, stage 4 of 14: decide

How do jobs reach workers?

A video is queued. Some process with CPU to spare has to find it, process it, and report back. Jobs take 5-20 minutes. You expect about 2,000 a day at launch and 10x within a year, which is still under one job every four seconds on average.

You have Postgres. A managed message queue (at-least-once delivery, visibility timeouts) is also available if you want one.

System so far· 5 parts
1234CLIENTInstructorbrowserSERVICEVideo APIDATABASEPostgresOBJECT STOREObject storageWORKERReconciler

Select a component to see what it is responsible for and which state it owns.

  1. 1Instructor browser → Video API: Create upload, report parts, poll status
  2. 2Instructor browser → Object storage: Upload parts via presigned URLs
  3. 3Video API → Object storage: Complete multipart upload, verify object
  4. 4Reconciler → Postgres: Find abandoned uploads and orphaned output
  • Request / response
  • Bulk data

What you need to know

0 of 2 checks done
  1. Work that outlives a request has to become a durable record somewhere: a row, or a message in a queue. A job that exists only in a process's memory dies with that process, and nothing remembers it was owed.

  2. When one action changes two systems, say "mark the video queued" in Postgres and "publish a job" to a message queue, a crash between the two leaves them disagreeing. That's the dual write problem.

    If both changes live in the same database, a single transaction makes them happen together or not at all.

  3. Think first

    The API updates the video to 'queued' in Postgres, then publishes a message to the queue. The process crashes between the two. What state is the system in?