TCP
TCP is defined in RFC 793, RFC 1122, RFC 2018, RFC 5681, RFC 7323.
The TCP connection
TCP is connection-oriented because two processes must first establish a connection via a handshake before exchanging data.
During this setup, they exchange preliminary segments to agree on connection parameters, and both sides initialize TCP state variables.
A TCP connection:
- consists of:
- buffers (send and receive)
- variables (state information, sequence numbers, congestion control state, timers, connection endpoints)
- sockets
- is point-to-point, between a single sender and a single receiver
- provides full-duplex communication, so data can flow simultaneously in both directions (A → B and B → A)
- does not support multicasting (transfer of data from one sender to many receivers in a single send operation)
Once a TCP connection is established:
- the client writes a stream of data to the socket
- TCP places the data in the connection's send buffer
- TCP periodically takes data from the send buffer and sends it to the network layer as segments
- segment size is limited by the Maximum Segment Size (MSS):
- the maximum application-layer data in a TCP segment
- derived from the link-layer Maximum Transmission Unit (MTU) and negotiated during the TCP handshake
- typical MTU (Ethernet, PPP) ≈ 1500 bytes
- typical MSS ≈ 1460 bytes
The network layer encapsulates TCP segments into IP datagrams and sends them across the network.
At the receiver:
- incoming segment data is stored in the TCP connection's receive buffer
- the application reads the stream of data from this buffer
TCP header
| Transmission Control Protocol (TCP) | ||||||
|---|---|---|---|---|---|---|
| Offsets (starting positions) |
Octet | 0 | 1 | 2 | 3 | |
| Octet | Bit | 0-3 | 4-7 | 8-15 | 16-23 | 24-31 |
| 0 | 0 | Source Port | Destination Port | |||
| 4 | 32 | Sequence Number | ||||
| 8 | 64 | Acknowledgment Number | ||||
| 12 | 96 | Header Length | Unused | Flags | Receive Window Size | |
| 16 | 128 | Checksum | Urgent Data Pointer | |||
| 20+ | 160+ | Options | ||||
Header fields:
| Name | Size | Description | |
|---|---|---|---|
Source port |
16 bits |
||
Destination port |
16 bits |
||
Sequence number |
32 bits |
Used for reliable data transfer. TCP treats data as a continuous stream of bytes:
In practice both sides of a TCP connection choose a pseudo-random Initial Sequence Number (ISN) to prevent sequence prediction attacks. |
|
Acknowledgment number |
32 bits |
Used for reliable data transfer. TCP uses cumulative acknowledgments, meaning it acknowledges everything up to the first missing byte:
What happens when segments arrive out of order?
|
|
Header length |
4 bits |
Length of the TCP header which can be of variable length due to the TCP options field. |
|
Flags |
CWR, ECE |
1 bit each |
Used in explicit congestion notification. |
URG |
1 bit |
Indicates that urgent data is present. |
|
ACK |
1 bit |
Indicates that the segment contains an acknowledgment for a segment that has been successfully received. |
|
PSH |
1 bit |
Indicates that the receiver should pass the data to the upper layer immediately. |
|
RST, SYN, FIN |
1 bit each |
Used for connection setup and teardown. |
|
Receive window size |
16 bits |
Used for flow control. |
|
Checksum |
16 bits |
||
Urgent data pointer |
16 bits |
Indicates the location of the last byte of this data marked as URG (rarely used). |
|
Options |
variable-length |
Optional. Used for features such as MSS, window scaling, timestamps, and SACK, typically negotiated during connection setup. |
|
RTT estimation and timeout
How large should TCP timeout intervals be?
They should be larger than the round-trip time (RTT) to avoid premature retransmissions.
TCP estimates RTT using three measurements:
| Sample (in ms) | Description | |
|---|---|---|
|
The time between sending a TCP segment and receiving its ACK. TCP does not use
|
|
|
Network conditions vary due to congestion in the routers and/or load on the end systems. So TCP smooths SampleRTT values using a running average called
For the first RTT measurement, TCP initializes |
|
|
The deviation in the RTT measures how far
Intuition:
|
|
Then, it uses these RTT measurements to estimate the timeout interval:
- timeout is set:
- above
EstimatedRTTto avoid unnecessary retransmissions - but not too much larger than
EstimatedRTT, otherwise:- when a segment is lost
- TCP would not quickly retransmit the segment
- leading to large data transfer delays
- above
- so TCP adds a safety margin to the
EstimatedRTTusingDevRTT, which is:- large when there is a lot of fluctuation in the
SampleRTTvalues - small when there is little fluctuation
- large when there is a lot of fluctuation in the
TimeoutInterval = EstimatedRTT + 4 * DevRTT- an initial
TimeoutIntervalvalue of 1 second is recommended (RFC 6298)
- an initial
- when a timeout occurs:
- TCP doubles
TimeoutInterval(exponential backoff) to reduce the chance of repeated premature retransmissions during congestion
- TCP doubles
- when an ACK provides a valid
SampleRTTmeasurement:- TCP updates
EstimatedRTTandDevRTT TimeoutIntervalis computed again using the formula
- TCP updates
Reliable data transfer
IP is unreliable:
- datagrams may be lost (e.g., buffers can overflow)
- datagrams may arrive out of order
- datagrams may be corrupted (bits can flip)
Because of this, IP only offers a "best effort" service.
TCP adds reliability on top of IP. It ensures that data received by an application is:
- uncorrupted
- without gaps
- without duplication
- in sequence (the byte stream is exactly the same byte stream that was sent)
TCP timer mechanism
TCP uses timers for reliability:
- associating a single timer for each segment would be too costly
- instead, TCP (per RFC 6298) uses a single retransmission timer (RTO) per connection
- this timer tracks the oldest unacknowledged data
If the timer expires, TCP retransmits unacknowledged data and adjusts its timeout estimate.
Simplified interesting scenarios:
| Scenario | Diagram |
|---|---|
|
One segment is sent:
|
|
|
Two segments are sent:
|
|
|
Two segments are sent:
|
|
Doubling the timeout interval
Most TCP implementations change how the timeout is calculated after a timeout occurs:
When TCP retransmits a packet:
- the next timeout is set to twice the previous timeout rather than deriving it from the last
EstimatedRTTandDevRTT - this causes the timeout interval to grow exponentially after repeated retransmissions
When there is no retransmission:
TimeoutIntervalis derived from the most recent values ofEstimatedRTTandDevRTT
This acts as a simple form of congestion control, because timeouts are usually caused by network congestion.
By waiting longer between retransmissions, TCP avoids making congestion worse.
Fast retransmit (retransmission before timeout)
Timeout-based retransmissions introduce delay because the sender detects a lost segment only after the retransmission timeout (RTO) expires.
To detect loss earlier, TCP uses duplicate ACKs:
- when the receiver detects a gap in the sequence numbers, it sends another ACK for the last in-order byte received
- after three duplicate ACKs, the sender assumes the missing segment was lost and retransmits it immediately, without waiting for the timeout
This mechanism is called fast retransmit.
Is TCP a GBN or a SR protocol?
Instead of acknowledging every segment it receives, the TCP receiver acknowledges the longest in-order sequence of data received.
This is similar to the acknowledgment strategy used in Go-Back-N (GBN).
However, TCP differs from GBN:
- TCP receivers buffer out-of-order segments
- if one ACK is lost, TCP usually retransmits only the missing segment, not everything after it
- example:
- segments
1, 3, 4are sent successfully but the ACK for segment2is lost - GBN retransmits segment
2and all later segments - TCP retransmits at most segment
2and may not retransmit anything if a later ACK arrives in time
- segments
With Selective Acknowledgment (SACK), TCP becomes even closer to Selective Repeat (SR):
- the receiver can explicitly report which out-of-order segments were received
- the sender retransmits only the missing data
Thus, TCP's error-recovery mechanism is best viewed as a hybrid of GBN and SR.
Flow control
TCP is full-duplex, so each host maintains a receive buffer for every connection (each socket has its own send and receive buffer):
- incoming TCP data is stored in the receive buffer
- the receiving application reads from this buffer
- if the receiving application reads slowly, the sender may overflow the buffer by sending too fast
To prevent this, TCP provides flow control, a speed-matching service between:
- rate at which the sender is sending
- rate at which the receiving buffer is read
This is different from congestion control, which slows the sender due to network congestion. Both throttle the sender, but for different reasons.
How TCP flow control works:
- each TCP receiver has a buffer of size
RcvBuffer - the free space in this buffer is the receive window
rwnd:- it changes dynamically as data is read
- initially,
rwnd = RcvBuffer
Example:
Host Asends data toHost BHost Badvertises its currentrwndin TCP segments sent back toHost A- this tells
Host Ahow much buffer space is available Host A:- tracks
LastByteSent – LastByteAcked = unacknowledged data in flight - ensures
unacknowledged data in flight ≤ rwndto guarantee the receiver buffer is never overflowed
- tracks
Zero window problem:
Host B's receive buffer becomes full and advertisesrwnd = 0Host Astops sending data- later,
Host B's application frees buffer space, butHost Bhas nothing more to send, soHost Ais not informed thatrwndhas increased
TCP solution:
Host Aperiodically sends small probe 1-byte segments whenrwnd = 0Host Backnowledges each probe- once buffer space becomes available, acknowledgments include a nonzero
rwnd, allowingHost Ato resume sending
In contrast, UDP does not provide flow control, so receiver buffers can overflow and data may be lost.
TCP connection life
During the life of a TCP connection, the TCP protocol running in each host makes transitions through various TCP states.
TCP connection establishment
TCP connection establishment can add delay because it requires a three-way handshake before a connection is fully established:
-
the client-side TCP sends a special
SYNsegment to request a connection with a server:SYN = 1(flag bit)Sequence number = client_isnisn= initial sequence number- chosen randomly to avoid certain security attacks (CERT 2001–09, RFC 4987)
-
upon receiving the
SYNsegment, the server-side TCP:- extracts the TCP
SYNsegment - allocates the TCP buffers and variables to the connection
- sends back a special
SYN-ACK"connection-granted" segment:SYN = 1(flag bit)ACK = client_isn + 1(flag bit)Sequence number = server_isn(the server chooses its own initial sequence number)
-
the "connection-granted" segment is referred to as a SYN-ACK segment:
I received your SYN packet to start a connection with your initial sequence number,
client_isn.
I agree to establish this connection.
My own initial sequence number isserver_isn
- extracts the TCP
-
upon receiving the
SYN-ACKsegment, the client-side TCP:- allocates buffers and variables to the connection
- sends the server another segment to acknowledge the server's connection-granted segment:
ACK = server_isn + 1SYN = 0(since the connection is established)
- this segment may carry application-layer data
This process uses three messages, so it's called the three-way handshake:
This process is also sometimes targeted by attacks such as SYN floods.
TCP connection closing
Either of the two processes participating in a TCP connection can end the connection.
When a connection ends, the "resources" (buffers and variables) in the hosts are deallocated.
If the client decides to close the connection:
- the client application process issues a close command
- this causes the client TCP to send a special TCP segment to the server process
- the
FINbit is set to1in the segment's header
- the
- when the server receives this segment
- it sends the client an acknowledgment segment in return
- then it sends its own shutdown segment
- the
FINbit is set to1in the segment's header
- the client acknowledges the server's shutdown segment
- at this point, all the resources in the two hosts are now deallocated
Reset segment
If a host receives a TCP SYN packet for port 443, but nothing is listening on that port (for example, no web server is running), it replies with a reset (RST) segment:
- the
RSTflag bit is set to 1 -
this tells the sender:
I don't have a socket for that segment. Please do not resend the segment.
How Nmap leverages TCP for port-scanning
To explore a specific TCP port (say port 6789) on a target host, nmap will send a TCP SYN segment with destination port 6789 to that host.
There are three possible outcomes for the source host:
- it receives a TCP
SYNACKsegment- this means an application is running with TCP port 6789
- nmap returns "open"
- it receives a TCP
RSTsegment- this means that the
SYNsegment reached the target host - but the target host is not running an application with TCP port 6789
- the attacker at least knows that the segments destined to the host at port 6789 are not blocked by any firewall on the path between source and target hosts
- this means that the
- it receives nothing
- this likely means that the
SYNsegment was blocked by an intervening firewall - and never reached the target host
- this likely means that the
Most the things nmap can do are done by manipulating TCP connection-management segments.