Design a Video Processing Pipeline, stage 10 of 14: break it
The video that kills every worker
The lease mechanism faithfully recovers the crashed job, again and again. Meanwhile other instructors' videos wait behind a job that will never succeed, and this instructor sees "processing" indefinitely.
System so far· 8 parts
Select a component to see what it is responsible for and which state it owns.
- 1Instructor browser → Video API: Create upload, report parts, poll status
- 2Instructor browser → Object storage: Upload parts via presigned URLs
- 3Video API → Object storage: Complete multipart upload, verify object
- 4Video API → Postgres: Video row and job row in one transaction
- 5Transcode workers → Postgres: Claim lease, heartbeat, fenced completion
- 6Transcode workers → Object storage: Read raw upload, write attempt output
- 7Reconciler → Postgres: Find abandoned uploads and orphaned output
- 8CDN → Object storage: Origin fetch on cache miss
- 9Student player → CDN: Manifest and segments
- Request / response
- Bulk data
What you need to know
Not every failure is worth retrying. Retries help only when the next attempt might behave differently:
Kind Example Retry? Transient storage returned 503, instance reclaimed yes, with backoff Permanent not a valid video, unsupported codec no: fail now, with a reason Unknown the worker died yes, but within a budget Check
A corrupt file crashes ffmpeg 40 seconds into every attempt. Leases recover the job each time. With no attempt limit, what happens?Failure needs to be visible to the person who can act on it. The video's state machine gets a terminal move,
processing → failed, with a reason written for the instructor, such as "the file appears to be corrupted; try re-exporting it", not "exit code 139". Engineers can still inspect the job record or a dead-letter table.Work it out
With a budget of 5 attempts, each crashing after 40 seconds, about how many worker-minutes does one poison file waste?