Skip to content

Design a Video Processing Pipeline, stage 2 of 14: decide

Where do the bytes go?

An instructor selects a 4 GB file. Your API servers are stateless, redeploy several times a day, and cut requests at 60 seconds. You have Postgres and an S3-compatible object store.

Choose where the file's bytes travel and where they come to rest.

System so far· 3 parts
1CLIENTInstructorbrowserSERVICEVideo APIDATABASEPostgres

Select a component to see what it is responsible for and which state it owns.

  1. 1Instructor browser → Video API: Create upload, report parts, poll status

What you need to know

0 of 2 checks done
  1. A system has two kinds of traffic:

    • The control plane: small, important decisions. Who may upload, which video this is, what state it's in. It needs authorization and belongs in your API and database.
    • The data plane: the bytes themselves. Here, gigabytes of video. It should take the shortest path to storage built for large blobs.

    Mixing them makes the API tier carry traffic it's bad at.

  2. Object storage (S3 and similar) stores files ("objects") by key, durably, at a low price per gigabyte. Two features matter here:

    • A presigned URL is a link your server signs that lets whoever holds it do one specific thing, such as upload part 3 of one object, until it expires. The client never sees storage credentials.
    • Multipart upload splits a file into parts uploaded separately, even in parallel. If one part fails, only that part is retried. See Object storage.
  3. Check

    Uploads stream through the API servers into storage. The API is redeployed 4 times a day. What happens to a 27-minute upload that's in progress during a deploy?