Design a Video Processing Pipeline, stage 8 of 14: break it
Two workers, one job
Leases recover crashed workers. But this worker did not crash. It only lost contact with the database for a while. Read the timeline and select the lines where the design (not the network) is at fault.
System so far· 8 parts
Select a component to see what it is responsible for and which state it owns.
- 1Instructor browser → Video API: Create upload, report parts, poll status
- 2Instructor browser → Object storage: Upload parts via presigned URLs
- 3Video API → Object storage: Complete multipart upload, verify object
- 4Video API → Postgres: Video row and job row in one transaction
- 5Transcode workers → Postgres: Claim lease, heartbeat, fenced completion
- 6Transcode workers → Object storage: Read raw upload, write attempt output
- 7Reconciler → Postgres: Find abandoned uploads and orphaned output
- 8CDN → Object storage: Origin fetch on cache miss
- 9Student player → CDN: Manifest and segments
- Request / response
- Bulk data
What you need to know
A lease can expire while its holder is still running. The holder may not know: it may be paused, or unable to reach the database. You can't prevent this, because from the outside a paused process looks exactly like a dead one.
What you can do is make the old holder's writes harmless.
Freeze worker A for longer than its lease, then switch on fencing. Each new lease comes with a higher token number, and storage remembers the highest token it has accepted.
A lease that runs out while its holder is asleep. The lease lasts 10 s and A renews it every 3 s, but A freezes at 2 s (a GC pause, a VM migration, a slow disk). Change how long A is frozen, and whether storage checks fencing tokens. 12 sWorker AWorker BStorage- 0 sWorker ATakes the lease with fencing token 33
- 2 sWorker AStalls (GC pause) for 12 s
- 10 sLockA's lease expires
- 10 sWorker BTakes the lease with fencing token 34
- 11 sWorker BWrites the job's result
- 11 sStorageAccepts B's write (token 34)
- 14 sWorker AResumes, still believes it holds the lease, and writes
- 14 sStorageAccepts A's write and overwrites B's result
Two workers acted as the owner. A's lease expired during the pause, B took over, and storage accepted A's late write anyway. The job's result is now whatever A wrote last.
Check
With fencing on, worker A resumes and writes with token 33 after worker B has written with token 34. What happens to A's write?In a database, fencing is a condition on the write:
UPDATE jobs SET status = 'succeeded' WHERE id = 812 AND lease_token = :mine;A stale worker's update matches zero rows. For files, which can't check tokens, give each attempt its own output location (for example
renditions/812/attempt-2/) so two attempts never write over each other.