← all topics

6 · ARP, ICMP, DHCP, DNS and the life of a packet Network

ARP request/reply · ICMP types · DHCP DORA and relay · DNS resolution chain · IPv6 NDP/SLAAC · three rehearsed 'what happens when' walk-throughs

Why it matters for NPE. 'Two hosts just booted, how do they start talking', 'ping across two networks' and 'type facebook.com into a browser' are the most-reported network questions after BGP and TCP. These four protocols are the moving parts.

Primer: four helper protocols, one story

A host that has just booted knows nothing. DHCP gives it an address, a mask, a gateway and a DNS server. DNS turns a name into an address. ARP (or IPv6 NDP) turns the next hop's IP into a MAC so a frame can be built. ICMP is how the network reports back when something goes wrong, and it is what ping and traceroute are made of. Every "what happens when" question is these four in order, then TCP, then the application.

Order to recite: DHCP (who am I) → DNS (who is the target) → route lookup (local or via gateway) → ARP (what MAC do I write) → send → ICMP tells you if it failed. Say which one you are in at every step.

Watch

Free CCNA | Ethernet LAN Switching (Part 2) | Day 6 | CCNA 200-301 Complete CourseJeremy's IT Lab · 33:41

The "first packet on a LAN" story with a capture. Watch 6:02-20:25: 6:02 ARP, 7:41 ARP request, 10:15 ARP reply, 12:04 ARP table, 15:31 ping, 18:28 Wireshark capture, 20:25 the MAC table filling as a side effect. Skip the quizzes from 27:06.

Free CCNA | DHCP | Day 39 | CCNA 200-301 Complete CourseJeremy's IT Lab · 37:02

Watch 10:15-21:35: 10:15 Discover, 12:45 Offer, 14:57 Request, 16:38 Ack, 17:58 DORA summary, 18:37 DHCP relay (why it matters in a routed network). Skip the Cisco config from 21:35.

Free CCNA | DNS | Day 38 | CCNA 200-301 Complete CourseJeremy's IT Lab · 30:11

Watch 2:02-11:09: 2:02 purpose, 4:40 nslookup, 6:31 a lookup in Wireshark, 9:36 DNS cache, 11:09 hosts file. Skip 12:15-20:24 (IOS DNS config).

How DNS Works - ComputerphileComputerphile · 8:04

Recursive resolver vs root, TLD and authoritative servers, and why caching makes the system work. Best short explanation of the hierarchy.

DHCP Explained - Dynamic Host Configuration ProtocolPowerCert Animated Videos · 10:09

Only if DORA has not stuck after Jeremy. Ten animated minutes.

Reading beats video for two things: the record-type table at Cloudflare, DNS records, and the IANA ICMP registry for the types and codes below.

ARP: IP on this link → MAC

ARP sits directly on Ethernet (EtherType 0x0806), has no IP header, and therefore never crosses a router. A host ARPs only for addresses in its own subnet or for its next hop.

ARP REQUEST (broadcast frame) ARP REPLY (unicast frame) Ethernet: dst ff:ff:ff:ff:ff:ff Ethernet: dst aa:aa:aa:aa:aa:aa (the asker) src aa:aa:aa:aa:aa:aa src bb:bb:bb:bb:bb:bb type 0x0806 type 0x0806 ARP: hardware type 1 (Ethernet) ARP: hardware type 1 protocol type 0x0800 (IPv4) protocol type 0x0800 hlen 6 plen 4 hlen 6 plen 4 opcode 1 (request) opcode 2 (reply) sender MAC aa:aa:aa:aa:aa:aa sender MAC bb:bb:bb:bb:bb:bb sender IP 10.1.1.10 sender IP 10.1.1.20 target MAC 00:00:00:00:00:00 (unknown) target MAC aa:aa:aa:aa:aa:aa target IP 10.1.1.20 target IP 10.1.1.10 "Who has 10.1.1.20? Tell 10.1.1.10" "10.1.1.20 is at bb:bb:bb:bb:bb:bb"
$ ip neigh show
10.1.1.1  dev eth0 lladdr 00:1c:73:aa:bb:01 REACHABLE
10.1.1.20 dev eth0 lladdr bb:bb:bb:bb:bb:bb STALE
10.1.1.99 dev eth0  FAILED                      <-- probed three times, nobody answered

$ tcpdump -eni eth0 arp
aa:aa:aa:aa:aa:aa > ff:ff:ff:ff:ff:ff, ethertype ARP (0x0806): Request who-has 10.1.1.20 tell 10.1.1.10
bb:bb:bb:bb:bb:bb > aa:aa:aa:aa:aa:aa, ethertype ARP (0x0806): Reply 10.1.1.20 is-at bb:bb:bb:bb:bb:bb

ICMP: the network's error channel

ICMP rides directly on IP (protocol 1), has no ports, and is generated by routers and hosts, not by applications. Error messages carry the IP header plus the first 8 bytes of the packet that caused them, so the sender can match the error to a socket.

TypeCodeMeaningWho sends it, when you see it
8 / 00Echo request / echo replyping. Identifier and sequence number in the header; payload usually carries a timestamp for RTT.
3 Destination unreachable0Network unreachableRouter has no route to that prefix.
1Host unreachableLast-hop router's ARP for the host failed.
2Protocol unreachableHost does not run that IP protocol number.
3Port unreachableHost has no UDP listener on that port. traceroute's finish line.
4Fragmentation needed and DF setPath MTU discovery. Block it and big transfers hang.
13Communication administratively prohibitedAn ACL or firewall configured to reject rather than drop.
11 Time exceeded0 / 1TTL hit zero in transit / fragment reassembly timed outEach router on the way sends code 0; traceroute is built on it.
5 Redirect1Redirect for hostGateway tells a host "use this other router on your subnet". Usually disabled on hosts for security.

DHCP: DORA

DHCP runs over UDP, client port 68, server port 67. The client has no address yet, so it sends from 0.0.0.0 to the broadcast address, and the switch floods it. Every message carries the same 32-bit transaction ID (xid) and the client's MAC (chaddr) so the client can match replies without an IP.

StepSrc IP:port → Dst IP:portDst MACWhat is inside
Discover (client)0.0.0.0:68 → 255.255.255.255:67ff:ff:ff:ff:ff:ffOption 53 type 1, chaddr, option 55 parameter request list (mask, router, DNS, domain), option 50 if it wants its old address back, option 61 client identifier.
Offer (server)10.1.1.1:67 → 255.255.255.255:68 (or unicast to the offered IP if the broadcast flag is clear)broadcast or client MACyiaddr = offered address, option 1 mask, 3 router, 6 DNS, 51 lease time, 54 server identifier.
Request (client)0.0.0.0:68 → 255.255.255.255:67ff:ff:ff:ff:ff:ffOption 50 requested IP and option 54 the chosen server. Broadcast so other servers that offered can withdraw.
Ack (server)10.1.1.1:67 → 255.255.255.255:68 (or unicast)broadcast or client MACSame options, lease starts. Client ARP-probes the address (duplicate check) and sends DHCPDECLINE if someone answers; a NAK from the server means start over.
$ tcpdump -ni eth0 -v udp port 67 or udp port 68
IP 0.0.0.0.68 > 255.255.255.255.67: BOOTP/DHCP, Request from aa:aa:aa:aa:aa:aa, xid 0x3f2a1c, DHCP-Message: Discover
IP 10.1.1.1.67 > 255.255.255.255.68: BOOTP/DHCP, Reply, xid 0x3f2a1c, Your-IP 10.1.1.10, DHCP-Message: Offer, Server-ID 10.1.1.1, Lease-Time 86400, Subnet-Mask 255.255.255.0, Default-Gateway 10.1.1.1, Domain-Name-Server 10.1.1.53
IP 0.0.0.0.68 > 255.255.255.255.67: BOOTP/DHCP, Request from aa:aa:aa:aa:aa:aa, xid 0x3f2a1c, DHCP-Message: Request, Requested-IP 10.1.1.10, Server-ID 10.1.1.1
IP 10.1.1.1.67 > 255.255.255.255.68: BOOTP/DHCP, Reply, xid 0x3f2a1c, Your-IP 10.1.1.10, DHCP-Message: ACK

DNS: name → record

browser → stub resolver (libc, /etc/hosts first, then /etc/resolv.conf nameserver) │ recursive query "A www.facebook.com?" (UDP 53) ▼ recursive resolver (10.1.1.53, ISP, 8.8.8.8 ...) ── cache hit? answer, done │ iterative queries, caches every answer it gets ├─▶ root server (a-m.root-servers.net, built in): "ask the .com servers" (NS + glue) ├─▶ .com TLD server: "ask ns1.facebook.com" (NS + glue) └─▶ authoritative facebook.com server: www.facebook.com CNAME star-mini.c10r.facebook.com star-mini.c10r.facebook.com A 157.240.x.x, TTL 60 ▼ answer back to stub, cached for TTL seconds at every level
RecordMapsInterview note
Aname → IPv4 addressOne name can return several A records (round-robin).
AAAAname → IPv6 addressClients ask for A and AAAA in parallel; Happy Eyeballs picks the faster.
CNAMEalias → canonical nameResolver follows the chain. A CNAME cannot sit at a zone apex alongside SOA/NS and cannot coexist with other records on the same name.
NSzone → authoritative server namesDelegation. "Glue" A records in the parent break the chicken-and-egg problem.
MXdomain → mail server name + preferenceLower preference wins. Points to a name, never an IP.
PTRIP → name (reverse)Under in-addr.arpa (10.1.1.20 is 20.1.1.10.in-addr.arpa) or ip6.arpa. dig -x.
SOAzone metadataPrimary server, serial, refresh/retry/expire, and the minimum TTL used for negative caching.
TXTname → free textSPF, DKIM, DMARC, domain ownership proofs.
SRV_service._proto.name → priority, weight, port, targetService discovery (LDAP, SIP, Kubernetes headless services).
$ dig www.facebook.com

;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 51283
;; flags: qr rd ra; QUERY: 1, ANSWER: 2, AUTHORITY: 0, ADDITIONAL: 1

;; QUESTION SECTION:
;www.facebook.com.              IN      A

;; ANSWER SECTION:
www.facebook.com.        3600   IN      CNAME   star-mini.c10r.facebook.com.
star-mini.c10r.facebook.com. 60 IN      A       157.240.22.35

;; Query time: 18 msec
;; SERVER: 10.1.1.53#53(10.1.1.53) (UDP)

Read it top down: status is the rcode. flags: qr rd ra means a response, recursion was asked for and is available; no aa because a cache answered, not the authority. The TTL column is what is left in the cache (60 s on the A record: Meta moves traffic fast). Query time near 0 ms is a cache hit; SERVER shows which resolver you actually used. Useful variants: dig @8.8.8.8 name to bypass the local resolver, dig +trace name to walk root → TLD → authoritative yourself, dig +short, dig -x 157.240.22.35, dig name AAAA.

IPv6 equivalents

IPv4IPv6Detail
ARP request (broadcast)Neighbor Solicitation, ICMPv6 type 135Sent to the solicited-node multicast ff02::1:ffXX:XXXX (last 24 bits of the target address), so only hosts with those bits process it. No broadcast exists in IPv6.
ARP replyNeighbor Advertisement, type 136Unicast back with the link-layer address option. Flags: S solicited, O override, R router.
Gratuitous ARPDuplicate Address DetectionNS for your own tentative address with source ::. An NA back means a conflict.
Default gateway from DHCP option 3Router Solicitation 133 to ff02::2, Router Advertisement 134 to ff02::1The RA carries the prefix (/64), the router's link-local as default gateway, MTU, and flags. DHCPv6 never hands out a gateway.
DHCPSLAAC or DHCPv6RA flag A: build your own address from the prefix (EUI-64 from the MAC, or a random/privacy interface ID). Flag M: get addresses from DHCPv6 (UDP 546 client / 547 server, to ff02::1:2, Solicit/Advertise/Request/Reply). Flag O: SLAAC for the address, DHCPv6 only for DNS and domain. RA option 25 (RDNSS) can carry DNS directly.
169.254/16fe80::/10 link-localEvery interface always has one, computed locally, no server needed. NDP and routing protocols run over it. Specify the interface to use it: ping fe80::1%eth0.
ICMP redirectICMPv6 Redirect, type 137Same intent.

ICMPv6 is not optional in IPv6: block types 133-137 and nothing on the link works; block type 2 (Packet Too Big) and PMTUD fails, and IPv6 routers never fragment. ip -6 neigh shows the neighbor cache with the same REACHABLE/STALE states as IPv4.

Three rehearsed walk-throughs

(a) Two hosts on one switch have just booted. How do they start communicating? And in IPv6?
  1. Link. Each NIC autonegotiates speed and duplex with the switch port. The switch MAC table is empty. Nothing has an IP yet.
  2. Address. Each host either has a static address or broadcasts a DHCP Discover from 0.0.0.0:68. The switch learns that host's MAC on the ingress port and floods the broadcast. If a DHCP server (or relay) is on the segment: Offer, Request, Ack, then the host ARP-probes its new address. If nothing answers after a few seconds, the host self-assigns 169.254.x.x.
  3. Decide. Host A (10.1.1.10/24) wants to reach 10.1.1.20. AND with the mask: same network, so the destination is local and A must resolve 10.1.1.20 directly, not the gateway.
  4. ARP. A has no entry, so it broadcasts "who has 10.1.1.20". The switch floods it. B matches the target IP, caches A's IP→MAC from the sender fields, and unicasts a reply. The switch learns B's MAC and forwards the reply out A's port only.
  5. Send. A builds the frame: dst MAC bb, src MAC aa, EtherType 0x0800, dst IP 10.1.1.20. The switch now knows both MACs and forwards without flooding. B replies the same way. If the application used a hostname, DNS (or mDNS on a server-less segment) happened before step 3.
  6. IPv6 variant. Each host computes a link-local fe80:: address from its MAC or a random ID, runs DAD (NS for its own address, source ::). It multicasts a Router Solicitation to ff02::2. With only two hosts and a switch there is no router, so no RA, no global prefix, no gateway: both hosts stay link-local, which is enough to talk. A multicasts a Neighbor Solicitation to B's solicited-node group ff02::1:ffXX:XXXX; the switch floods it (or restricts it with MLD snooping). B answers with a unicast Neighbor Advertisement. A sends to B's link-local; the application must pass the interface (%eth0). If a router were present, its RA would add a /64 and SLAAC addresses in the same second.
(b) Host pings a host in another network through switch, router, switch. What changes at each hop?
  1. A (10.1.1.10/24, gw 10.1.1.1) → S1. A ANDs 10.2.2.30 with its mask: different network, so look up the route table, default via 10.1.1.1, ARP for 10.1.1.1 (not for 10.2.2.30). Frame: dst MAC r1a, src aa; packet: src 10.1.1.10, dst 10.2.2.30, TTL 64, protocol 1; ICMP type 8, id, seq 1.
  2. S1. Reads only the dst MAC, looks it up, forwards out the router's port (floods if unknown). Learns aa on A's port. Changes nothing in the frame.
  3. R1. Accepts the frame because dst MAC is its own, strips Ethernet, checks TTL > 1, longest-prefix match on 10.2.2.30 gives a connected route out the other interface, decrements TTL to 63, recomputes the IP checksum, ARPs for 10.2.2.30 on that interface (first time) and builds a new frame: src r1b, dst bb. IPs unchanged. If the ARP fails, R1 sends A an ICMP type 3 code 1 host unreachable. If there were no route, type 3 code 0.
  4. S2. Floods the router's ARP request, learns r1b, forwards the echo request to B's port once it knows bb.
  5. B. dst MAC mine, dst IP mine, protocol 1, type 8: build type 0 echo reply with the same id, seq and payload. B's own decision: 10.1.1.10 is remote, ARP for its gateway 10.2.2.1, send to r1b.
  6. Return. R1 routes back, TTL 63, new frame to aa. A matches id and seq and prints the RTT. Say out loud: MACs changed at the router, IPs never did, TTL dropped by one, and the reply needed its own route and its own ARP on both sides. Then say where it fails: no default route on A (A gives "network unreachable" locally), no return route on R1 or B's gateway, ACL dropping ICMP (silent) or rejecting (type 3 code 13), B's host firewall ignoring echo.
(c) You type https://www.facebook.com in a browser. What happens?
  1. Address (DHCP). Already done at boot: the host has an IP, mask, gateway and resolver. Failure: 169.254 address, nothing further works.
  2. Name (DNS). Browser cache, then OS cache and /etc/hosts, then the stub sends an A and an AAAA query over UDP 53 to the configured resolver. The resolver answers from cache or walks root → .com → facebook.com authoritative. Answer: CNAME star-mini.c10r.facebook.com, then A (and AAAA) with a short TTL. Failures: SERVFAIL (resolver cannot reach authority), NXDOMAIN (typo), timeout (UDP 53 blocked or resolver down). Symptom: "server not found" instantly or after a few seconds, while curl https://157.240.22.35 still works.
  3. Path (route + ARP). 157.240.22.35 is remote: default route, ARP for the gateway MAC (cache hit normally). Failure: no gateway entry, FAILED neighbor, wrong VLAN. Symptom: everything beyond the subnet dead, local hosts fine.
  4. Transport (TCP). SYN to port 443 with MSS, window scale, SACK; SYN+ACK; ACK. One RTT. Failure: SYN retransmitted with no reply (filtered or no route back), RST (nothing listening). Modern browsers also race QUIC over UDP 443 and use it if the server advertised HTTP/3 via Alt-Svc.
  5. Security (TLS 1.3). ClientHello with SNI www.facebook.com, supported ciphers, key share; ServerHello, certificate chain, Finished; client verifies the chain against its trust store and checks the name, sends Finished. One more RTT, then everything is encrypted. Failures: certificate expired or name mismatch (browser warning), clock wrong, middlebox stripping TLS, protocol version mismatch. Symptom: TCP connects, then an immediate close or an alert.
  6. Application (HTTP). GET / HTTP/2 with Host, cookies, Accept headers. 200 OK with HTML, or 301/302 to another host, which restarts at step 2. The HTML references dozens of assets on static CDN names, each needing DNS and often a connection; HTTP/2 multiplexes them over one TCP connection, HTTP/3 over QUIC. Failures: 5xx from the server, slow responses with everything below healthy (an application problem, not a network one).
  7. Finish strong. Name the evidence for each stage: dig, ip route get, ip neigh, ss -tn or tcpdump for the handshake, openssl s_client -connect host:443 -servername host for TLS, curl -v for HTTP. That is also the troubleshooting answer in topic 19.

Interview questions

1. What is ARP and where does the ARP table live?ARP maps an IP address on the local link to a MAC address. The cache lives on every IP host and router, per interface. A broadcast request asks "who has this IP", the owner replies unicast, and both sides cache the mapping. Switches do not have an ARP table for forwarding; they have a MAC table, which is MAC to port. Follow-up: what happens when it expires? Linux marks it STALE, keeps using it, and re-probes on the next send; if three unicast probes fail, the entry is FAILED and traffic drops.
2. Why does a host never ARP for a remote address?ARP only works on the local link because it has no IP header and routers do not forward it. For a remote destination the host ARPs for its next hop, the gateway, and puts the gateway's MAC with the remote IP in the same frame.
3. What is gratuitous ARP and when would you want it?An unsolicited broadcast announcing your own IP-to-MAC mapping. Hosts send it at boot to detect duplicates, and a VIP owner sends it on failover so neighbors update their caches immediately. Without it the old MAC stays cached until the entry times out, which on a Cisco router is four hours.
4. Walk me through DORA with the addresses and ports.Discover from 0.0.0.0 port 68 to 255.255.255.255 port 67, broadcast MAC, carrying the client MAC and a transaction ID. Offer from the server's IP port 67 to the broadcast address port 68 with yiaddr, mask, router, DNS, lease and server ID. Request from 0.0.0.0 again, broadcast, naming the requested IP and the chosen server so others withdraw. Ack confirms and the lease starts. Follow-up: in a routed network the gateway relays the broadcast as a unicast with its address in giaddr, and the server picks the scope from that.
5. A host has 169.254.12.7. What does that tell you?It asked for DHCP and nothing answered, so it self-assigned a link-local address. The host and its link are up; the problem is between it and the DHCP server: wrong VLAN on the port, the VLAN missing from a trunk, no relay on the gateway, scope exhausted, or the server down. Check whether other hosts on the same VLAN get leases to split the problem.
6. When does a DHCP client renew, and what if the server is gone?At T1, half the lease, it unicasts a Request to its server. If no reply, at T2, 87.5%, it broadcasts a Request so any server can renew it. At expiry it must drop the address and start with Discover. So a dead DHCP server does not cause an immediate outage; it causes one spread over the lease time, which is a classic "it broke gradually" story.
7. Explain DNS resolution end to end.The application asks the stub resolver, which checks /etc/hosts and its cache, then sends a recursive query to the configured resolver. The resolver, if not cached, asks a root server, which refers it to the TLD servers, which refer it to the zone's authoritative servers, which answer. Each answer is cached for its TTL. Follow-up: the stub does one recursive query; the resolver does iterative queries on its behalf.
8. NXDOMAIN or SERVFAIL: which tells you what?NXDOMAIN is an authoritative answer that the name does not exist, so the DNS system worked and the name or zone is wrong. SERVFAIL means the resolver failed to get an answer: the authoritative servers were unreachable, the delegation is broken, or DNSSEC validation failed. SERVFAIL points at infrastructure; NXDOMAIN points at data.
9. Is DNS UDP or TCP?Both. Queries go over UDP 53 by default. If the response exceeds the UDP size (512 bytes classic, larger with EDNS0), the server sets TC and the client retries over TCP 53. Zone transfers are always TCP, and DoT (853) and DoH (443) are TCP too. A firewall allowing only UDP 53 breaks big answers.
10. How does traceroute work and what do the stars mean?It sends probes with increasing TTL. Each router that decrements TTL to zero returns ICMP Time Exceeded, type 11, from its ingress address. The destination returns Port Unreachable (UDP probes) or Echo Reply (ICMP probes), ending the trace. Stars at one hop with later hops answering mean that router rate-limits or filters ICMP. Stars through to the end mean the path or the return path breaks there.
11. What replaces ARP and DHCP in IPv6?Neighbor Discovery over ICMPv6. Neighbor Solicitation to the solicited-node multicast and Neighbor Advertisement replace ARP. Router Advertisements give the prefix and default gateway, and with the A flag hosts build their own address (SLAAC). DHCPv6 is optional, for addresses with the M flag or just DNS with the O flag, and it never provides the gateway. Every interface also has a link-local fe80:: address without any server.
12. Which ICMP types would you never block at a firewall?Type 3 code 4, fragmentation needed, or PMTUD breaks and large transfers hang. Type 11 so traceroute works. For IPv6, types 1-4 (errors, including Packet Too Big) and 133-137 (NDP), because without NDP nothing on the link resolves. Echo can be rate-limited but blocking it blinds your own monitoring.

Traps

Scenario

A user says "the internal wiki is down". You find curl https://wiki.corp.example hangs for five seconds then says "could not resolve host", but curl -k https://10.20.30.40 returns the wiki login page. Other users are fine. Walk through it.

Expected reasoning
  1. Scope. Works by IP, not by name, from one host: the path, TCP and the application are healthy. This is name resolution on this host or between this host and its resolver. Say that first.
  2. What resolver is it using? cat /etc/resolv.conf (or resolvectl status). A VPN client or a manual change may have left a resolver that is unreachable from this network. Compare with a working host.
  3. Can it reach that resolver? dig @10.1.1.53 wiki.corp.example. A timeout means UDP 53 to the resolver is blocked, the resolver is down, or there is no route. Try dig @8.8.8.8 example.com: if that works, the host's DNS transport is fine and the problem is the specific resolver.
  4. What does the resolver say? If it answers, read the status. NXDOMAIN: the host is appending the wrong search domain or the record was removed; try the FQDN with a trailing dot and dig +search. SERVFAIL: the resolver cannot reach the internal authoritative server, which would affect everyone, so recheck whether others really are fine. NOERROR with an answer: the resolver is fine and the stub is the problem.
  5. Stub side. getent hosts wiki.corp.example uses the same path the application does. Check /etc/hosts for a stale override, /etc/nsswitch.conf order, and a local caching daemon (systemd-resolved, nscd, dnsmasq) that may hold a stale negative answer: resolvectl flush-caches.
  6. Confirm and close. Fix the resolver config or flush the cache, rerun curl -v and watch "Trying 10.20.30.40:443" appear. State the root cause and the evidence in one sentence. If it had failed by IP as well, you would have gone down the route, ARP, TCP, TLS ladder instead.

← 5 · File handling and log parsing · all topics · 7 · Sorting, Big-O and top-k with heaps →