Design a Video Processing Pipeline, stage 11 of 14: break it
Deleting a video mid-transcode
Deletion touches every store in the system: the video row, the job row, raw and rendered objects, and whatever the CDN has cached. A worker may be mid-flight. Evaluate each statement about this situation.
System so far· 8 parts
Select a component to see what it is responsible for and which state it owns.
- 1Instructor browser → Video API: Create upload, report parts, poll status
- 2Instructor browser → Object storage: Upload parts via presigned URLs
- 3Video API → Object storage: Complete multipart upload, verify object
- 4Video API → Postgres: Video row and job row in one transaction
- 5Transcode workers → Postgres: Claim lease, heartbeat, fenced completion
- 6Transcode workers → Object storage: Read raw upload, write attempt output
- 7Reconciler → Postgres: Find abandoned uploads and orphaned output
- 8CDN → Object storage: Origin fetch on cache miss
- 9Student player → CDN: Manifest and segments
- Request / response
- Bulk data
What you need to know
Deleting a video touches several stores with no shared transaction: the database rows, objects in storage, and copies cached at CDN edges. A worker may also be in the middle of a job for it.
A tombstone makes deletion a state rather than an absence:
status = 'deleted'. Anything later that's conditioned on another status (like the worker'sWHERE status = 'processing') quietly fails to apply.Check
Which order is safe if the process can crash at any step?Two kinds of leftovers, and they're not equally bad:
- Garbage: data nothing references. Costs money. Safe to clean up later.
- Dangling references: a reference to data that's gone. Users see broken behaviour.
Order distributed steps so that every crash leaves garbage, never dangling references.
Think first
The original objects are deleted from storage. A student loads the video's manifest URL through the CDN a minute later. What might they get?