TCP congestion control

The congestion-control mechanism is a key component of TCP:

End-to-end congestion

Congestion control helps TCP figure out how much data the network can handle, so the sender knows how many packets can be sent without overwhelming the network.

The sender adjusts its sending rate based on network congestion:

This leads to three main questions.

How does a sender limit the rate at which it sends data?

TCP controls the sending rate using a variable called Congestion Window or cwnd:

The size of cwnd is dynamic, function of perceived network congestion.

The sender limits how much unacknowledged data can be in transit:

TCP cwd

How does a sender detect congestion?

TCP assumes the network is congested when a loss event happens. This can be:

These events suggest that packets were lost because of congestion on the path between sender and receiver.

If no loss events occur, TCP assumes the network is working well:

Because TCP uses incoming ACKs to control when it increases cwnd, TCP is said to be self-clocking.

How should the sender adjust its sending rate?

TCP senders must collectively balance their sending rate:

TCP adjusts its sending rate using these rules:

TCP also continuously probes the available bandwidth to adjust the transmission rate:

Additive increase/multiplicative decrease (AIMD)

TCP works on many kinds of networks, from slow links to extremely fast ones. So there is no way for a TCP sender to know the network's capacity.

In the late 1980s, Van Jacobson introduced the first TCP congestion-control algorithm in RFC 5681. This happened about eight years after TCP/IP was already in use, when the Internet was experiencing congestion collapse.

The algorithm controls the congestion window (cwnd) using additive increase/multiplicative decrease (AIMD):

So TCP decreases cwnd aggressively and increases cwnd conservatively throughout the lifetime of the connection.

This creates a sawtooth pattern:

TCP additive increase/multiplicative decrease (AIMD)

Intuition:

The algorithm has three main parts:

  1. slow start
  2. congestion avoidance
  3. fast recovery

Slow start

Additive increase works well when TCP is already near the network's capacity.

But when a connection starts from scratch, TCP needs a faster way to increase its sending rate. Ironically, this mechanism is called slow start, which increases cwnd exponentially rather than linearly.

During slow start, the cwnd grows like this:

TCP effectively doubles the number of packets it has in transit every RTT:

TCP slow start

Slow start ends in three cases:

  1. timeout (loss detected):
    • the TCP sender assumes congestion
    • cwnd is reset to 1 MSS
    • a new threshold, ssthresh, is set to half of the old cwnd
    • a second state variable ssthresh (slow start threshold) is set to half of the value of cwnd when congestion was detected
    • TCP restarts slow start
  2. reaching the slow start threshold (cwnd >= ssthresh)
    • TCP becomes more cautious:
      • it might be reckless to keep doubling cwnd when it reaches ssthresh
      • congestion could be just around the corner
    • slow start ends
    • TCP switches to congestion avoidance mode and increases cwnd more cautiously
  3. three duplicate ACKs detected before a timeout:
    • TCP assumes a segment was lost but the network is still usable
    • it retransmits the missing segment (fast-retransmit)
    • and enters the fast recovery mode

Congestion avoidance

When TCP enters congestion avoidance:

When does congestion avoidance stop?

  1. timeout occurs:
    • TCP sets ssthresh = cwnd / 2
    • TCP sets cwnd = 1 MSS
    • TCP switches back to slow start
  2. three duplicate ACKs detected before a timeout:
    • TCP retransmits the missing segment (fast-retransmit)
    • TCP enters the fast recovery

Fast retransmit and fast recovery

Fast recovery is an optional TCP feature.

Normally, TCP retransmits lost packets only after a timeout. This can led to long periods of time during which the connection is dead while waiting for a timer to expire.

Fast retransmit resends a dropped packet after receiving 3 duplicate ACKs, instead of waiting for a timeout:

This reduces the number of timeouts and improves throughput (often by around 20%).

After fast retransmit, TCP enters fast recovery:

Slow start is used only at the beginning of a connection or after a timeout.

At all other times, cwnd is following a pure additive increase/multiplicative decrease pattern.

TCP CUBIC

Additive-increase/multiplicative-decrease (AIMD) congestion control creates a "sawtooth" pattern.

It illustrates the intuition of TCP "probing" for available bandwidth.

But is cutting the rate in half and then increasing it slowly the best strategy?

This idea is the basis of TCP CUBIC (RFC 8312):

TCP CUBIC is the default TCP congestion control algorithm in Linux.

Other congestion-avoidance approaches

Loss-based algorithms in TCP Reno/Tahoe interpret packet loss as a signal of congestion.

Other congestion-avoidance approaches try to proactively detect congestion before packet loss occurs.

Network-assisted congestion control

Explicit Congestion Notification (ECN) (RFC 3168) involves both TCP and IP to signal network congestion.

At the IP layer, ECN uses two bits (4 possible values) in the IP header.

These bits are used in two ways:

  1. routers mark packets to indicate congestion:
    • congestion indication is carried to the destination host
    • the destination then informs the sending host using TCP feedback
    • the definition of when a router is congested is a configuration choice decided by the network operator
  2. hosts uses them to indicate support for ECN:
    • this allows routers to mark packets instead of dropping them when congestion occurs
    • ECN capability is negotiated during TCP connection setup

Delay-based congestion control

Delay- or Avoidance-based algorithms try to detect congestion before packets are dropped, using measurements such as throughput, RTT, and bottleneck bandwidth.

Keep the pipe just full, but no fuller.

TCP Vegas (1994):

Variants of Vegas:

TCP BBR (Bottleneck Bandwidth and RTT):

Fairness

A congestion-control mechanism is said to be fair if all connections get an equal share of bandwidth on a link.

AIMD (additive-increase/multiplicative-decrease) tends to promote fairness:

UDP traffic does not behave fairly:

Also, nothing stops a TCP-based application from using multiple parallel connections:

Previous Congestion control All ⏎ Next Introduction to the network layer (data plane)

A Kemar Joint