Video Compression Explained: Codecs, Containers, Bitrate & Keyframes
Demystifying how video compression works under the hood. Understand H.264 vs HEVC vs AV1, GOP structures, I/P/B frames, and CRF vs CBR rate control.
The Miracle of Video Compression
An uncompressed 4K video (3840x2160 pixels) at 60 frames per second with 10-bit color depth generates approximately **11.9 Gigabits of raw data every single second**. A standard two-hour feature film in raw format would require more than **10 Terabytes** of storage.
Yet on YouTube, Netflix, or your local phone camera, that same two-hour film fits comfortably inside a 2 GB to 6 GB file. This 500:1 compression ratio is not magic—it is the result of four decades of perceptual mathematics and spatial-temporal redundancy reduction.
Containers vs. Codecs: The Essential Distinction
One of the most frequent misconceptions among digital creators is confusing a **container** with a **codec**:
- **The Container (.mp4, .mkv, .webm, .mov)**: Think of the container as an envelope or box. It bundles the video stream, audio stream, subtitles, chapter markers, and color space metadata into a single synchronized file.
- **The Codec (H.264/AVC, H.265/HEVC, AV1, VP9, ProRes)**: The codec is the algorithm used to mathematically compress and decompress the pixels themselves.
For example, an `.mp4` file can house an older H.264 video track, a modern AV1 video track, or even a HEVC 10-bit HDR stream.
The Modern Codec Landscape Compared
| Codec | Standard | Efficiency vs H.264 | Hardware Decode Support | Royalties & Patents | Best Use Case |
|---|---|---|---|---|---|
| H.264 (AVC) | 2003 | Baseline (1.0x) | 99.9% universal on all CPUs/GPUs | MPEG-LA patent pool | Universal web fallback & legacy players |
| H.265 (HEVC) | 2013 | ~50% better | Broad (iOS, modern PCs, 4K TVs) | Fragmented patent pools (Velos, HEVC Advance) | 4K HDR recording on iPhone, drone cameras |
| VP9 | 2013 | ~45% better | All modern browsers & Android | Open source / Royalty-free (Google) | YouTube 1440p and 4K web streaming |
| AV1 (AOMedia) | 2018 | ~65% better | Apple M3+, Intel 11th Gen+, RTX 40+ | Open source / Royalty-free consortium | The future of high-efficiency web video |
How Temporal Compression Works: I, P, and B Frames
Raw video possesses two kinds of redundancy: 1. **Spatial Redundancy (Intra-frame)**: Pixels next to each other in the same frame (like a blue sky or flat wall) share nearly identical colors. 2. **Temporal Redundancy (Inter-frame)**: From one frame to the next (1/60th of a second), 95% of the scene remains unchanged; only a person's hand or mouth moves.
Codecs exploit temporal redundancy through a structure known as a **GOP (Group of Pictures)**, built from three frame types:
- **I-Frames (Keyframes)**: Completely self-contained images. They compress like standard JPEG photos without referencing any other frame. Scrubbing video player sliders jumps directly to the nearest I-Frame.
- **P-Frames (Predicted Frames)**: Store only what has changed relative to the previous frame using motion vectors.
- **B-Frames (Bi-directional Frames)**: Look both backward to previous frames and forward to upcoming frames to calculate motion interpolations, achieving maximum compression.
Rate Control Methods: CRF vs. CBR vs. VBR
How should your encoder allocate bits?