TCP

TCP is defined in RFC 793, RFC 1122, RFC 2018, RFC 5681, RFC 7323.

The TCP connection

TCP is connection-oriented because two processes must first establish a connection via a handshake before exchanging data.

During this setup, they exchange preliminary segments to agree on connection parameters, and both sides initialize TCP state variables.

A TCP connection:

TCP connection

Once a TCP connection is established:

The network layer encapsulates TCP segments into IP datagrams and sends them across the network.

At the receiver:

TCP header

Transmission Control Protocol (TCP)
Offsets
(starting positions)
Octet 0 1 2 3
Octet Bit 0-3 4-7 8-15 16-23 24-31
0 0 Source Port Destination Port
4 32 Sequence Number
8 64 Acknowledgment Number
12 96 Header Length Unused Flags Receive Window Size
16 128 Checksum Urgent Data Pointer
20+ 160+ Options

Header fields:

Name Size Description

Source port

16 bits

Destination port

16 bits

Sequence number

32 bits

Used for reliable data transfer.

TCP treats data as a continuous stream of bytes:

  • the segments are simply containers for carrying pieces of the byte stream
  • sequence numbers refer to byte positions in this stream
Byte stream:

    SEQ:        0               1000            2000             3000
                |---------------|---------------|---------------|
    bytes:      ABCDEFGHI       JKLMNOPQR       STUVWXYZ        ...

TCP segments:

    Segment 1
    SEQ=0
    bytes=ABCDEFGHI
    carries bytes 0–999

                    Segment 2
                    SEQ=1000
                    bytes=JKLMNOPQR
                    carries bytes 1000–1999

                                Segment 3
                                SEQ=2000
                                bytes=STUVWXYZ
                                carries bytes 2000–2999

In practice both sides of a TCP connection choose a pseudo-random Initial Sequence Number (ISN) to prevent sequence prediction attacks.

Acknowledgment number

32 bits

Used for reliable data transfer.

TCP uses cumulative acknowledgments, meaning it acknowledges everything up to the first missing byte:

  • if a host receives:
    • bytes 0–535
    • bytes 900–1000
    • but is missing 536–899
  • it will acknowledge 536, since that is the next expected byte

What happens when segments arrive out of order?

  • the TCP RFCs do not impose any rules here and leave the decision up to the TCP implementation
  • choices:
    1. discard them until missing data arrives (less common today)
    2. buffer them and fill gaps later (common)
    3. use Selective Acknowledgment (SACK) when available

Header length

4 bits

Length of the TCP header which can be of variable length due to the TCP options field.

Flags

CWR, ECE

1 bit each

Used in explicit congestion notification.

URG

1 bit

Indicates that urgent data is present.

ACK

1 bit

Indicates that the segment contains an acknowledgment for a segment that has been successfully received.

PSH

1 bit

Indicates that the receiver should pass the data to the upper layer immediately.

RST, SYN, FIN

1 bit each

Used for connection setup and teardown.

Receive window size

16 bits

Used for flow control.

Checksum

16 bits

Urgent data pointer

16 bits

Indicates the location of the last byte of this data marked as URG (rarely used).

Options

variable-length

Optional.

Used for features such as MSS, window scaling, timestamps, and SACK, typically negotiated during connection setup.

RTT estimation and timeout

How large should TCP timeout intervals be?

They should be larger than the round-trip time (RTT) to avoid premature retransmissions.

TCP estimates RTT using three measurements:

Sample (in ms) Description

SampleRTT

The time between sending a TCP segment and receiving its ACK.

TCP does not use SampleRTT measurements for retransmitted segments because it cannot know whether the ACK acknowledges the original transmission or the retransmission, making the RTT measurement ambiguous:

Time 0 ms:    send segment (original)
Time 100 ms:  retransmit segment
Time 150 ms:  receive ACK

Measured time:
send → ACK = 150 ms

Possible RTT:
- original transmission: 150 ms
- retransmission: 50 ms

EstimatedRTT

Network conditions vary due to congestion in the routers and/or load on the end systems.

So TCP smooths SampleRTT values using a running average called EstimatedRTT:

  • EstimatedRTT = (1- α) * EstimatedRTT + α * SampleRTT
  • where α = 1/8 = 0.125 (RFC 6298)

EstimatedRTT is an exponential weighted moving average (EWMA):

  • recent measurements have more influence than older ones
  • this is natural because recent values better reflect current network conditions

For the first RTT measurement, TCP initializes EstimatedRTT with SampleRTT.

SampleRTT and EstimatedRTT

DevRTT

The deviation in the RTT measures how far SampleRTT deviates from EstimatedRTT.

  • DevRTT = (1- β) * DevRTT + β * |SampleRTT - EstimatedRTT|
  • where β = 1/4 = 0.25 (RFC 6298)

Intuition:

  • if SampleRTT values have little fluctuation → DevRTT will be small
  • if there is a lot of fluctuation → DevRTT will be large

Then, it uses these RTT measurements to estimate the timeout interval:

Reliable data transfer

IP is unreliable:

Because of this, IP only offers a "best effort" service.

TCP adds reliability on top of IP. It ensures that data received by an application is:

TCP timer mechanism

TCP uses timers for reliability:

If the timer expires, TCP retransmits unacknowledged data and adjusts its timeout estimate.

Simplified interesting scenarios:

Scenario Diagram

One segment is sent:

  • its ACK is lost
  • Host A retransmits after a timeout
  • Host B recognizes that the sequence number corresponds to data already received
  • Host B discards the retransmitted data and resends the ACK
Retransmission due to a lost acknowledgment

Two segments are sent:

  • both arrive correctly at Host B, which sends cumulative ACKs
  • the ACKs are delayed and do not reach Host A before the timeout
  • Host A assumes loss and retransmits only the oldest unacknowledged segment, because TCP assumes later data may already have been received and buffered
  • Host A restarts its timer
  • before the new timeout expires, the delayed ACKs finally arrive at Host A, acknowledging both segments
  • since the second segment is now acknowledged, it is not retransmitted
Segment 120 not retransmited

Two segments are sent:

  • one ACK is lost, but a later cumulative ACK confirms all data was received
  • Host A does not retransmit anything
A cumulative ACK avoids retransmission of the first segment

Doubling the timeout interval

Most TCP implementations change how the timeout is calculated after a timeout occurs:

When TCP retransmits a packet:

When there is no retransmission:

This acts as a simple form of congestion control, because timeouts are usually caused by network congestion.

By waiting longer between retransmissions, TCP avoids making congestion worse.

Fast retransmit (retransmission before timeout)

Timeout-based retransmissions introduce delay because the sender detects a lost segment only after the retransmission timeout (RTO) expires.

To detect loss earlier, TCP uses duplicate ACKs:

This mechanism is called fast retransmit.

Is TCP a GBN or a SR protocol?

Instead of acknowledging every segment it receives, the TCP receiver acknowledges the longest in-order sequence of data received.

This is similar to the acknowledgment strategy used in Go-Back-N (GBN).

However, TCP differs from GBN:

With Selective Acknowledgment (SACK), TCP becomes even closer to Selective Repeat (SR):

Thus, TCP's error-recovery mechanism is best viewed as a hybrid of GBN and SR.

Flow control

TCP is full-duplex, so each host maintains a receive buffer for every connection (each socket has its own send and receive buffer):

To prevent this, TCP provides flow control, a speed-matching service between:

This is different from congestion control, which slows the sender due to network congestion. Both throttle the sender, but for different reasons.

How TCP flow control works:

TCP receive buffer

Example:

Zero window problem:

TCP solution:

In contrast, UDP does not provide flow control, so receiver buffers can overflow and data may be lost.

TCP connection life

During the life of a TCP connection, the TCP protocol running in each host makes transitions through various TCP states.

TCP connection establishment

TCP connection establishment can add delay because it requires a three-way handshake before a connection is fully established:

  1. the client-side TCP sends a special SYN segment to request a connection with a server:

    1. SYN = 1 (flag bit)
    2. Sequence number = client_isn
  2. upon receiving the SYN segment, the server-side TCP:

    • extracts the TCP SYN segment
    • allocates the TCP buffers and variables to the connection
    • sends back a special SYN-ACK "connection-granted" segment:
      1. SYN = 1 (flag bit)
      2. ACK = client_isn + 1 (flag bit)
      3. Sequence number = server_isn (the server chooses its own initial sequence number)
    • the "connection-granted" segment is referred to as a SYN-ACK segment:

      I received your SYN packet to start a connection with your initial sequence number, client_isn.
      I agree to establish this connection.
      My own initial sequence number is server_isn

  3. upon receiving the SYN-ACK segment, the client-side TCP:

    • allocates buffers and variables to the connection
    • sends the server another segment to acknowledge the server's connection-granted segment:
      • ACK = server_isn + 1
      • SYN = 0 (since the connection is established)
    • this segment may carry application-layer data

This process uses three messages, so it's called the three-way handshake:

TCP connection establishment

This process is also sometimes targeted by attacks such as SYN floods.

TCP connection closing

Either of the two processes participating in a TCP connection can end the connection.

When a connection ends, the "resources" (buffers and variables) in the hosts are deallocated.

If the client decides to close the connection:

TCP connection teardown

Reset segment

If a host receives a TCP SYN packet for port 443, but nothing is listening on that port (for example, no web server is running), it replies with a reset (RST) segment:

How Nmap leverages TCP for port-scanning

To explore a specific TCP port (say port 6789) on a target host, nmap will send a TCP SYN segment with destination port 6789 to that host.

There are three possible outcomes for the source host:

  1. it receives a TCP SYNACK segment
    • this means an application is running with TCP port 6789
    • nmap returns "open"
  2. it receives a TCP RST segment
    • this means that the SYN segment reached the target host
    • but the target host is not running an application with TCP port 6789
    • the attacker at least knows that the segments destined to the host at port 6789 are not blocked by any firewall on the path between source and target hosts
  3. it receives nothing
    • this likely means that the SYN segment was blocked by an intervening firewall
    • and never reached the target host

Most the things nmap can do are done by manipulating TCP connection-management segments.

Previous Reliable data transfer All ⏎ Next Congestion control

A Kemar Joint