Video streaming and CDNs
Streaming video (including Netflix, YouTube and Amazon Prime) account for about 55%–80% of Internet traffic, depending on the measurement method.
Video streaming services are implemented using:
- application-level protocols
- geographically distributed servers that function in some ways like a cache
Internet video
Prerecorded videos:
- are placed on servers
- users send requests to the servers to view the videos on demand
A video is:
- a sequence of images displayed at a constant rate (e.g., 24, 30 or 60 images per second)
- an uncompressed image consists of an array of pixels
- with each pixel encoded into a number of bits to represent luminance and color
Video can be compressed:
- trading off video quality with bit rate
- the higher the bit rate, the better the image quality
Compressed Internet video ranges from:
| Bitrate | Resolution |
|---|---|
| 300 kbps – 2 Mbps | 240p–480p |
| 2 – 5 Mbps | 720p HD |
| 4 – 10 Mbps | 1080p Full HD |
| 15 – 40 Mbps | 4K UHD |
| 50+ Mbps | 8K |
To provide continuous playout, the network must provide an average throughput to the streaming application that is at least as large as the bit rate of the compressed video.
Streaming systems use codecs such as H.264/AVC, HEVC/H.265 and AV1.
HTTP streaming
HTTP streaming:
- the video is stored at an HTTP server as an ordinary file or collection of media segments with a specific URL
- the client:
- establishes a TCP connection with the server
- issues HTTP requests for the video content
- the server:
- sends the video data
- within HTTP response messages
- as quickly as the underlying network protocols and traffic conditions will allow
- bytes are collected in a client application buffer
- once the number of bytes in this buffer exceeds a predetermined threshold
- the client application begins play-back
- the streaming video application periodically grabs video frames from the client application buffer
- once the number of bytes in this buffer exceeds a predetermined threshold
If all clients receive the same encoding of the video, users with low bandwidth may experience stalls while users with high available bandwidth may not fully utilize the network.
This has led to the development of adaptive HTTP-based streaming protocols:
- Dynamic Adaptive Streaming over HTTP (DASH)
- HTTP Live Streaming (HLS)
DASH:
- the video is encoded into several different versions
- the HTTP server has a manifest file:
- provides a URL for each version along with its bit rate
- the client first requests the manifest file and learns about the various versions
- the client:
- dynamically requests chunks or media segments of a few seconds in length
- typically requests segments using URLs listed in the manifest file
- may also use HTTP byte-range requests in some deployments
- adapts to bandwidth changes during the session by measuring the received bandwidth and running a rate determination algorithm to select the next chunk to request:
- chunks from a high-rate version when the amount of available bandwidth is high
- chunks from a low-rate version when the available bandwidth is low
HLS:
- developed by Apple
- provides adaptive bitrate streaming using
M3U8playlist files and segmented media - is widely used on Apple devices and platforms
CMAF is a universal shipping container for internet video streaming, so that DASH and HLS can share the same media segments.
Content Distribution Networks
Major video-streaming companies use Content Distribution Networks (CDNs).
A CDN consists of a set of geographically distributed servers that cache content and serve user requests from locations that are close to the user or otherwise optimal.
CDNs use routing mechanisms such as DNS-based redirection, Anycast, and load balancing to direct users to an appropriate server.
CDNs have two different server placement philosophies:
- enter deep:
- place many small CDN server clusters inside access ISPs, close to end users
- this reduces latency and improves user experience by minimizing network distance
- but makes the system harder to manage due to its large number of sites
- bring home:
- use fewer, larger CDN clusters at centralized locations such as IXPs or major backbone sites
- this simplifies management
- but results in higher latency compared to "enter deep" deployments, depending on user location and network topology
In practice, CDNs use a hybrid approach, combining both strategies.
Netflix
Netflix video distribution has two major components:
- the Amazon cloud infrastructure
- the web site and its databases run in cloud infrastructure
- many different formats and bit rates are created for each movie
- its own private CDN infrastructure (Netflix Open Connect)
- Netflix has installed server racks both in IXPs and within residential ISPs themselves
- encoded versions of videos are uploaded to CDN servers
Netflix distributes by pushing videos to CDN servers during off-peak hours.
For locations that cannot hold the entire library, Netflix pushes only the most popular videos, which are determined on a day-to-day basis.
When a user selects a movie to play:
- the Netflix software:
- first determines which CDN servers have copies of the movie
- then determines an appropriate server for that client request
- sends the client the IP address of the selected server along with a manifest file
- the client and the CDN server then directly interact using adaptive HTTP streaming
- the client requests chunks from different versions of the movie
- chunks are typically a few seconds long
- the client measures the received throughput
- and runs a rate-determination algorithm to determine the quality of the next chunk to request
YouTube
The Google/YouTube design and protocols are proprietary.
Through several independent measurement efforts we can gain a basic understanding about how YouTube operates.
YouTube:
- makes extensive use of its own private CDN
- has installed server clusters in many hundreds of different IXP and ISP locations
- uses pull caching and DNS redirect
- directs clients using factors such as latency, server load, topology and cache availability
- sometimes a client is directed to a more distant cluster to balance load across clusters
- employs HTTP adaptive streaming
- uses HTTP byte-range requests and segmented streaming to control the amount of prefetched video data
- increasingly uses HTTP/3 and QUIC transport