π― The takeaway, first
Ingest once, transcode many, serve from the edge. The upload path and the watch path are two almost completely separate systems that meet at a blob store. The origin is boring on purpose β all the engineering lives in the transcode ladder (one upload becomes 1080p/720p/480p/360pβ¦) and the CDN (those copies live near viewers, in 6-second segments). If your design has viewers downloading the uploaded file, start over.
π« Misconception, busted
"YouTube stores one file per video and streams it." There is no "the file." Every upload becomes a ladder of renditions, each chopped into ~6-second segments, described by a manifest (HLS .m3u8 / DASH .mpd). The player downloads segments one at a time and switches renditions mid-video as your bandwidth changes. A 10-minute video is ~100 separate HTTP requests, not one stream.
Requirements
Scope it before you draw it.
Functional
- Upload video (resumable β uploads fail halfway)
- Transcode to multiple qualities + generate thumbnails
- Playback with adaptive bitrate
- Video metadata: title, description, duration, status
- View counts, likes (mention; counters are their own subsystem)
Non-functional
- 500 hours uploaded/min; 1B hours watched/day
- Start playback in < 2s; rebuffer ratio < 1%
- 11 nines durability on masters; 99.9% serving availability
- Works on a 1 Mbps connection and a 100 Mbps one
- Upload must survive a dead phone battery mid-transfer
Back-of-the-envelope math
Assumptions labeled. The egress row is the one that matters.
| What | Assumption | Math |
|---|---|---|
| Raw ingest | 500 hr/min, avg upload 5 Mbps β 0.625 MB/s | 720K hr/day Γ 3600s Γ 0.625 MB β 1.6 EB/day raw |
| Transcode ladder | 4 renditions β Γ2.5 total bytes | β 4 EB/day new stored β the ladder multiplies everything |
| Transcode compute | 720K content-hr/day, 4 renditions, ~0.5Γ realtime each | need β 60K concurrent transcode workers just to keep up |
| Egress (the bill) | 1B hrs watched/day, avg 720p @ 2.5 Mbps = 0.3125 MB/s | 3.6B s Γ 0.3125 MB β 1.1 EB/day β 104 Tbps average |
| CDN offload | 95% served from edge | origin only sees β 5 Tbps β this is why the CDN is the system |
| Hot set | 90% of views hit <10% of videos | edge cache is small; the long tail lives in cheap cold storage |
Go deeper: reality-checking the 4 EB/day
Real ingest is lower: most uploads are phone-bitrate 720p, per-title encoding shrinks easy content (a talking head compresses far better than confetti), and exact-duplicate uploads get deduped. The interview point isn't the exact number β it's that the ladder multiplies storage and compute, so every efficiency win (better codec, per-title bitrates) pays off N times. Say that sentence out loud; it's worth points.
Architecture
Two pipelines that meet at the origin store.
flowchart LR
U[Uploader] --> UA[Upload API
resumable sessions]
UA --> RAW[(Raw Blob Store
masters)]
RAW --> TQ[Transcode Queue]
TQ --> TW[Transcode Workers
FFmpeg fleet]
TW --> TH[Thumbnail Service]
TW --> OR[(Origin Store
renditions + segments + manifests)]
TH --> OR
OR --> CDN[CDN Edge PoPs
worldwide cache]
CDN --> P[Player]
UA --> MDB[(Metadata DB
Postgres)]
TW --> MDB
P -->|view events| AN[Analytics
async]
Uploads land in the raw store and get a 202 Accepted β transcoding is async. Workers pull jobs, produce the ladder plus thumbnails into the origin store, and flip the video status to ready. Viewers never touch the origin directly: the player fetches the manifest and segments from the nearest edge PoP, which backfills from origin on miss.
Go deeper: why is the player in charge of quality?
Adaptive bitrate is client-driven. The server just hosts segments and a manifest listing the options. The player measures its own throughput, keeps a few seconds of buffer, and picks the next segment's rendition. This is beautifully scalable: no per-viewer state on the server, and quality decisions happen where the information (actual bandwidth) lives. The server's only job is to make switching cheap β hence short segments and aligned keyframes.
Component deep-dives
Upload β resumable, because networks lie
sequenceDiagram
participant C as Uploader
participant A as Upload API
participant S as Raw Blob Store
participant Q as Transcode Queue
C->>A: POST /v1/uploads {name, size} β {upload_id}
loop chunks (e.g. 8 MB)
C->>A: PUT chunk (offset) β retryable, idempotent
A->>S: append bytes
end
C->>A: POST /v1/uploads/{id}/complete
A->>S: verify checksum
A->>Q: enqueue transcode job
A-->>C: 202 Accepted {video_id, status: processing}
Playback β the adaptive loop
sequenceDiagram
participant P as Player
participant E as CDN Edge
participant O as Origin
P->>E: GET /v/{id}/manifest.m3u8
E-->>P: rendition list (1080p/720p/480p/360p)
loop every ~6s segment
P->>P: measure throughput, check buffer
P->>E: GET segment (chosen rendition)
alt edge miss
E->>O: backfill segment
O-->>E: bytes
end
E-->>P: segment bytes
Note over P: buffer low? step down.
buffer healthy 5s? step up.
end
API + data model
POST /v1/uploads { "name": "...", "size": 1234567890 } β { "upload_id": "u_..." }
PUT /v1/uploads/{id}/chunk?offset=8388608 (idempotent, retryable)
POST /v1/uploads/{id}/complete β 202 { "video_id": "v_...", "status": "processing" }
GET /v1/videos/{id} β metadata + status
GET /v1/videos/{id}/manifest.m3u8 β HLS manifest (served from CDN)
videos(video_id PK, owner_id, title, duration_s, status, created_at) renditions(video_id, quality, bitrate_kbps, codec, segment_count) -- PK (video_id, quality) upload_sessions(upload_id PK, video_id, size, bytes_received, checksum, expires_at)
Trade-offs
| Decision | Option A | Option B | Pick |
|---|---|---|---|
| Manifest format | HLS (Apple everywhere) | DASH (more flexible) | Both β serve from the same segments via CMAF |
| Encoding | Fixed ladder (simple) | Per-title (bitrate matched to content complexity) | Fixed first, per-title when the compute bill hurts |
| Transcode timing | Eager (all renditions at upload) | Lazy (transcode on first watch) | Eager for popular creators, lazy for the long tail |
| Segment length | Shorter (faster quality switches) | Longer (better compression, fewer requests) | ~6s β the industry compromise |
Failure modes
What I'd actually build
Opinionated. Steal this for the interview.
Stack: TUS-protocol resumable uploads into S3 (raw bucket, versioned). SQS β Fargate Spot FFmpeg workers producing an HLS/CMAF ladder (1080p/720p/480p/360p) + WebP thumbnails. Origin = S3 + CloudFront with an origin-shield tier. Postgres for metadata. Player: hls.js with a custom ABR rule (throughput + buffer, hysteresis on step-up). Progressive publishing: 360p goes live first so creators see their video in under a minute. Everything behind signed URLs/cookies for private videos.
Interview tips
- Lead with the egress math. "104 Tbps average" justifies the entire CDN discussion that follows.
- Split the diagram in two early. Upload pipeline vs serving pipeline β interviewers watch for this separation.
- Say "the client picks the quality." ABR being client-driven surprises people and shows you know how streaming actually works.
- Resumable uploads are a free point. Mention chunked, idempotent, checksummed uploads before they ask about failures.
- Close with the long tail. Lazy transcode + cold storage for unwatched videos shows cost awareness.
π¬ Interactive: upload β transcode β CDN
Upload a video and watch it become a ladder of renditions, then play network-weather-god with the adaptive bitrate lab.
1 Β· Ingest waiting
Resumable: 8 MB chunks, idempotent retries, checksum on complete.
2 Β· Transcode ladder waiting
FFmpeg fleet: each rendition is chopped into ~6s segments as it encodes. 360p finishes first β progressive publishing.
3 Β· CDN propagate waiting
0 / 28 edge PoPs warm. Viewers fetch from the nearest green dot β the origin only gets asked on a miss.
4 Β· Adaptive bitrate lab live
You are the network. Drag the slider β or let it storm on its own.
Rules: step down instantly when bandwidth can't sustain the rendition; step up only after 5s of healthy buffer (hysteresis) β otherwise the picture yo-yos.