8 · TCP vs UDP: handshake, state, loss and congestion control Network
3-way handshake, seq/ack, FIN vs RST, retransmission, flow vs congestion control, MTU/MSS/PMTUD
Why it matters for NPE. TCP is the protocol most candidates are asked to explain in depth. Expect follow-ups on what you'd see in a capture when things go wrong.
Primer: TCP gives you a reliable byte stream; UDP gives you a datagram and a shrug
TCP numbers every byte, acknowledges what arrived, retransmits what didn't, keeps data in order, and slows down when the receiver or the network can't keep up. UDP adds only ports and a checksum to IP; anything else (ordering, retries) is the application's job, which is exactly why DNS, DHCP, QUIC and most real-time traffic use it. The interviewer wants to hear that you know which fields and timers make TCP's promises true and what you'd see on the wire when they fail.
Watch
The best single TCP explanation on YouTube. Key minutes: 1:32 seq/ack count bytes, 6:49 retransmission timer, 9:48 cumulative and delayed ACKs, 12:15 window and bytes in flight, 21:50 duplicate ACKs in troubleshooting, 24:50 the handshake, 31:22 FIN vs RST. Stop at 42:00.
Flow control vs congestion control, asked directly in interviews. 1:58 receive window vs send window, 6:31 how cwnd grows, 12:57 rebuilding cwnd after loss.
Depth with real captures. Watch 17:58-1:00:22: handshake, receive window, options (MSS, SACK, window scale, timestamps), window scaling.
MSS = MTU − 40, advertised per direction. Seven minutes, worth it.
RTO vs fast retransmit in a capture, with a low-MTU root cause at 6:12.
Watch 4:18-8:57 (ports, multiplexing) and 18:34-25:22 (UDP, the comparison table, port numbers to know).
See the handshake in a capture once so "what would you see in tcpdump" has a picture behind it.
Reading: RFC 9293 §3.5 (state diagram) and RFC 5681 §3.1-3.2 (slow start, congestion avoidance, fast retransmit).
Headers
| Field | What it does | Interview angle |
|---|---|---|
| Ports | Identify the socket. A connection is the 4-tuple (src IP, src port, dst IP, dst port). | Why can thousands of clients hit port 443? Because the client side of the tuple differs. |
| Sequence number | Byte offset of the first data byte in this segment. SYN and FIN each consume one sequence number. | "ACK 1001" means "I have everything up to byte 1000, send 1001 next". |
| ACK number | Next byte expected. ACKs are cumulative. | Delayed ACKs: receivers may ACK every other segment or after 40-200 ms. |
| Window | How many bytes the receiver can accept right now (flow control). 16 bits, so the window-scale option multiplies it. | Window 0 = "stop". Zero-window probes keep it alive. |
| Flags | SYN open, ACK, FIN polite close, RST abort, PSH deliver now. | RST means "no socket / go away"; a firewall reject vs a silent drop. |
| MSS option | Largest payload each side will accept, advertised in the SYNs. Usually MTU − 40 = 1460. | MSS clamping fixes PMTUD black holes. |
| SACK option | Selective ACK: "I have 1001-2000 and 3001-4000", so only the gap is resent. | Without SACK, one loss can trigger resending everything after it. |
Connection lifecycle
- Why three messages? Both sides must prove they can send and receive and must exchange their random initial sequence numbers. Two would leave the server unsure the client received its SYN+ACK.
- Why random ISNs? So a stale segment from an old connection on the same 4-tuple isn't mistaken for new data, and so off-path attackers can't guess sequence numbers.
- Half-open and SYN floods. The server keeps state after the SYN+ACK; floods exhaust the backlog. SYN cookies encode the state in the ISN instead.
- TIME_WAIT lets late segments die and makes sure the final ACK can be resent. Lots of TIME_WAIT sockets on a busy client is normal, not a leak.
- FIN vs RST. FIN is a polite "I have no more data", each direction closes independently (half-close). RST aborts immediately: sent when a segment arrives for a port with no listener, or when an application closes with unread data, or by a firewall that rejects rather than drops.
Reliability: timers and duplicate ACKs
- Retransmission timeout (RTO) is computed from smoothed RTT and its variance (RFC 6298: RTO = SRTT + 4×RTTVAR, min 200 ms on Linux, initial 1 s). On timeout the segment is resent and the RTO doubles (exponential backoff).
- Fast retransmit. If the receiver gets an out-of-order segment it re-ACKs the last in-order byte. Three duplicate ACKs tell the sender a segment was lost before the timer fires. With SACK it knows exactly which one.
- Out of order arrival is buffered by the receiver and acknowledged cumulatively once the gap fills. Reordering looks like loss to the sender if it is severe (spurious retransmits).
Flow control vs congestion control
| Flow control | Congestion control | |
|---|---|---|
| Protects | the receiver's buffer | the network (queues in switches and routers) |
| Signal | receive window (rwnd) advertised in every ACK | loss (timeout, dup ACKs), ECN marks, or delay (BBR) |
| Sender limit | bytes in flight ≤ min(rwnd, cwnd) | |
| Algorithm | sliding window | slow start (cwnd doubles per RTT) → congestion avoidance (+1 MSS per RTT) → on loss: halve (Reno/CUBIC) or back to 1 MSS on timeout. Linux default is CUBIC; BBR is model-based. |
Throughput ≈ window / RTT. A 64 KB window over a 100 ms RTT gives at most 5.2 Mbit/s regardless of link speed. That is the answer to "the link is 10 Gbit/s, why is one transfer slow across the ocean": the bandwidth-delay product exceeds the window, or loss keeps cwnd small (Mathis: throughput ≈ MSS / (RTT × √loss)).
MTU, MSS and the black hole
Ethernet MTU 1500 → IPv4 TCP MSS 1460. If a path has a smaller MTU (a VPN or MPLS tunnel, say 1400) and the packet has DF set, the router drops it and sends ICMP Fragmentation Needed (type 3 code 4). The sender then shrinks its segments: that is Path MTU Discovery. If a firewall blocks that ICMP, large segments vanish silently: the handshake works, small requests work, big responses hang. Fixes: allow ICMP type 3/4, lower the MTU, or clamp MSS on the edge router (--clamp-mss-to-pmtu). IPv6 routers never fragment; PMTUD is mandatory.
TCP vs UDP: when to choose which
| TCP | UDP |
|---|---|
| HTTP/1-2, SSH, BGP (port 179), SMTP, database connections, file transfer | DNS (53, with TCP fallback for big answers and zone transfers), DHCP (67/68), NTP (123), SNMP (161), syslog (514), VoIP/video, QUIC/HTTP3 (443), VXLAN (4789) |
| Need every byte, in order, exactly once | Need low latency; a late packet is useless; or the app has its own reliability (QUIC) or the exchange is one request/one reply |
| Cost: handshake RTT, head-of-line blocking, per-connection state | Cost: you handle loss, ordering, congestion yourself; easier to spoof and amplify |
What you'd see in a capture or in ss
$ ss -tni state established '( dport = :443 )'
ESTAB 0 0 10.1.1.10:51234 203.0.113.5:443 cubic wscale:7,7 rto:204 rtt:3.1/0.5 mss:1448 cwnd:10 bytes_sent:4210 retrans:0/2 ...
$ tcpdump -ni eth0 'tcp port 443 and (tcp[tcpflags] & (tcp-syn|tcp-rst) != 0)'
IP 10.1.1.10.51234 > 203.0.113.5.443: Flags [S], seq 1234, win 64240, options [mss 1460,sackOK,wscale 7]
IP 203.0.113.5.443 > 10.1.1.10.51234: Flags [S.], seq 5678, ack 1235, win 65160, options [mss 1460,sackOK,wscale 7]
IP 10.1.1.10.51234 > 203.0.113.5.443: Flags [.], ack 1, win 502
| Symptom in a capture | Likely meaning |
|---|---|
| SYN, SYN, SYN with backoff, no SYN+ACK | No route, silent drop by a firewall, or the host is down |
| SYN then RST+ACK immediately | Host reachable, nothing listening on that port (or firewall "reject") |
| Handshake OK, then retransmissions and dup ACKs | Loss on the path; check interface errors and queue drops hop by hop |
| Window 0 from one side | That application isn't reading its socket (slow consumer), not a network problem |
| Handshake OK, small transfers fine, big ones stall | PMTUD black hole |
| Many connections in SYN_RECV on the server | SYN flood or asymmetric path where the ACK never returns |
Interview questions
1. Walk me through the three-way handshake and what each side learns.
Client sends SYN with its ISN and options (MSS, window scale, SACK). Server replies SYN+ACK with its own ISN and ack = client ISN + 1. Client sends ACK with ack = server ISN + 1. After this both sides know the other can send and receive, both ISNs are agreed and the options are negotiated. Data can ride on the final ACK.2. Why not two messages?
The server would never know its SYN+ACK arrived, so it could not safely allocate the connection, and old duplicate SYNs could open bogus connections.3. What is the difference between flow control and congestion control?
Flow control protects the receiver using the advertised window. Congestion control protects the network using the congestion window the sender infers from loss or delay. The sender obeys the smaller of the two.4. How does TCP detect loss?
Two ways: the retransmission timer expires (slow, RTO ≥ 200 ms on Linux, doubling each time), or it receives three duplicate ACKs, which triggers fast retransmit without waiting for the timer. SACK tells it exactly which bytes are missing.5. What happens when packets arrive out of order?
The receiver buffers them and keeps ACKing the last in-order byte (duplicate ACKs). When the gap fills, it sends one cumulative ACK for everything. Heavy reordering makes the sender think there is loss and cut its window.6. What is TIME_WAIT and is it a problem?
The side that closes first waits 2×MSL (60 s on Linux) so late segments die and the last ACK can be retransmitted. Thousands of them on a busy client are normal; they only matter if you exhaust ephemeral ports to one destination, which you fix with connection reuse, not by disabling TIME_WAIT.7. Explain MSS vs MTU.
MTU is the largest IP packet a link carries (1500 on Ethernet). MSS is the largest TCP payload, MTU minus IP and TCP headers (1460). MSS is advertised in the SYN; MTU is a link property. A path with a smaller MTU in the middle needs PMTUD or MSS clamping.8. Why does one TCP flow not fill a 10 Gbit/s link over 100 ms RTT?
Throughput is bounded by window / RTT. To fill 10 Gbit/s at 100 ms you need 125 MB in flight; default windows and any loss (cwnd halving) keep it far below. Fixes: larger buffers and window scaling, BBR, parallel streams, or moving the data closer.9. When would you choose UDP?
When latency beats completeness (voice, video, gaming), when the exchange is a single request/reply (DNS, DHCP, NTP), or when the application implements its own reliability and congestion control (QUIC). Mention the cost: no backpressure, amplification risk.10. SYN goes out, no SYN+ACK comes back. What are the possibilities?
No route from client to server; no return route from server to client (asymmetric); a firewall silently dropping; the server down or its NIC/ARP broken; or the server received it and replied but the reply was dropped on the way back. Capture on both ends to split the problem in half.11. What is a RST and when do you see one?
A reset aborts the connection immediately. Sent when a segment hits a closed port, when a host receives a segment for a connection it doesn't have (after a reboot), when an app closes with unread data, or by a firewall configured to reject. Seeing SYN → RST means "reachable but not listening".12. What does the receive window being 0 tell you?
The receiving application is not draining its socket buffer. The sender must pause and probe. This is an application or host CPU problem, not a network problem.Traps
- Saying ACK numbers count packets. They count bytes; the ACK is the next expected byte.
- Saying "TCP is slower because it's reliable" without naming the mechanisms (handshake RTT, head-of-line blocking, cwnd ramp).
- Saying the server sends FIN "to close the connection". Each direction closes separately; CLOSE_WAIT piling up means your app never called close().
- Claiming UDP has no checksum. It does (optional in IPv4, mandatory in IPv6).
- Confusing the receive window (flow control, advertised) with the congestion window (sender's private estimate).
- Forgetting that DNS also uses TCP (answers over 512/1232 bytes, zone transfers), so "DNS is UDP" is incomplete.
- Calling BGP a "layer 3 protocol". It is an application that runs over TCP port 179.
Scenario
Users in California say a service in Europe is "slow". Pings show 150 ms RTT and 0.5% loss. The web team says the server is fine. What do you say and do?
Expected reasoning
- Translate the numbers: 0.5% loss at 150 ms RTT caps a single CUBIC flow at roughly MSS/(RTT×√p) ≈ 1460×8/(0.15×0.07) ≈ 1.1 Mbit/s. Loss, not latency, is the throughput killer here.
- Reproduce with a capture on the client: are there retransmissions and dup ACKs? Does cwnd (
ss -ti) stay small? That confirms loss on the path rather than a slow server. - Localise the loss with
mtrtoward the server and from the server back (asymmetric paths); look for the first hop where loss persists through to the end, and ignore hops that only rate-limit ICMP. - Check interface counters (errors, drops, discards) on the suspect hop, link utilisation and any policer/QoS. A dirty optic or a congested peering link are the usual causes.
- Mitigations while the path is fixed: move users to a closer edge/POP, enable BBR on the server, or use parallel connections. Report back with the evidence, not a guess.
← 7 · Sorting, Big-O and top-k with heaps · all topics · 9 · Strings, two pointers and sliding windows →