← all topics

19 · Troubleshooting walkthroughs Network

'Server unreachable' · 'latency and packet loss' · 'the site is slow' · layered method · DNS→route→ARP→TCP→TLS→HTTP · NAT and firewalls as definitions

Why it matters for NPE. 8 of 22 reports include a scenario. Candidates who pass describe scoping the problem, testing one layer at a time and naming the evidence each command gives. That calm structure is what is graded.

Primer: the method is what gets graded

Eight of the twenty-two reports behind this site include a troubleshooting scenario. Nobody failed for not knowing a command. People failed for jumping to a cause, for testing one direction only, and for not saying what each result would mean before running it. The interviewer is watching for a calm, layered, evidence-driven walk: scope it, split it, test one layer, keep what you learn, say where you are.

Say this skeleton out loud at the start of every scenario: "First I scope: who, what, since when, what changed. Then I split the path in half with a test on each side. Then one layer at a time, each test able to prove me wrong. I keep the evidence and I report status as I go." Then do it.

Watch

What happens when you type google.com into your browser and press enter? (Detailed Analysis)Hussein Nasser · 45:02

The long-form version of the most common opener, at the depth a 45-minute interview reaches. Watch 3:00 to 33:38: 7:40 DNS lookup and caches, 19:15 TCP connection, 24:19 TLS with SNI and ALPN, 33:38 the first GET. Stop before HTML parsing.

What happens when you type a URL into your browser?ByteByteGo · 5:20

The two-minute script. Memorise this ordering as your opening answer, then go deep where the interviewer steers.

Network Troubleshooting Methodology - CompTIA Network+ N10-009 - 5.1Professor Messer · 8:03

Identify, theorise, test, plan, implement, verify, document. Name these steps explicitly when handed a scenario.

TLS Handshake - EVERYTHING that happens when you visit an HTTPS websitePractical Networking · 27:58

The TLS leg. 3:12 Client Hello, 7:58 certificate chain, 11:36 keys, 26:13 TLS 1.3 changes (one round trip, no RSA key exchange).

Free CCNA | NAT (Part 1) | Day 44 | CCNA 200-301 Complete CourseJeremy's IT Lab · 32:10

Watch 1:34 to 15:03: private ranges, why NAT exists, static NAT. Skip the Cisco config after that; PAT is defined below.

Top 5 Wireshark tricks to troubleshoot SLOW networksDavid Bombal · 42:59

"The site is slow" with packets. Watch 7:59 to 38:58: handshake RTT, 21:42 finding slow packets with delta time, 25:37 dup ACKs, retransmits and zero window, 34:56 root cause.

Reading: Alex Gaynor's what-happens-when repo (the source Hussein follows); Julia Evans' networking notes; RFC 1918 §3 for the private ranges.

The method

  1. Scope. Who is affected: one user, one site, one region, everyone? What exactly fails: the error text, the port, by name or by IP? Since when, to the minute? What changed: a deploy, a firewall push, a maintenance, a new link, a DNS edit? Most incidents are solved by the answer to "what changed".
  2. Split the path in half. Pick a point in the middle and test from there in both directions. A capture at the client and at the server at the same moment tells you which half holds the fault. Then split that half.
  3. One layer at a time. Bottom-up when the host is suspect (link, address, ARP, route, firewall, TCP, TLS, HTTP). Top-down when the application is suspect (does curl work locally, does DNS resolve, does TCP connect). Say which direction you chose and why.
  4. Hypothesis, then a test that can falsify it. "I think the firewall drops it. If so, tcpdump on the server will show the SYN arrive and no SYN-ACK leave, and the DROP counter will move. If the SYN never arrives, I am wrong and the fault is upstream." Say what each outcome would mean before you look.
  5. Keep the evidence. Save the mtr, the pcap, the dig output. You will need them for the handoff and the postmortem.
  6. Change one thing at a time and re-test. Two changes at once means you cannot tell which one worked, and one of them may have made something else worse.
  7. Communicate. Every few minutes: "Confirmed X. Ruled out Y. Testing Z next. No ETA yet." Mitigate first (drain, fail over, roll back) and root-cause second when users are down.
client ---- switch ---- router ---- WAN ---- router ---- switch ---- server ^ ^ tcpdump here: did it leave? tcpdump here: did it arrive? did a reply leave? left, arrived, reply left, reply arrived → network fine, look at the application left, never arrived → fault between, split again arrived, no reply → server: firewall, no listener, rp_filter arrived, reply left, reply never arrived → return path (asymmetric routing)

Layered checklist

LayerQuestionCommandFail looks like
1 PhysicalLink up, right speed, clean?ip -br link, ethtool eth0, ethtool -S eth0NO-CARRIER, 1G on a 10G port, CRC errors climbing
2 LinkCan I resolve the next hop? Right VLAN?ip neigh, arping -I eth0, switch MAC tableFAILED/INCOMPLETE, two MACs for one IP, MAC learned on the wrong port
3 AddressRight IP and mask? Same subnet decision correct?ip addr/25 on one side and /24 on the other, 169.254.x.x (no DHCP)
3 RouteWhich way out, and is there a way back?ip route get X, traceroute -n, same from the far end"Network is unreachable", a stray /32, replies stop at hop N, forward and reverse paths differ
3 ReachabilityDoes IP work end to end?ping -c 20, mtr -rwzc 100Silence, Destination Host Unreachable, loss that persists to the last hop
3/4 FilterIs a firewall dropping it?nft list ruleset, iptables -L -n -v, conntrack -S, ACL/security groupDROP counter moving, SYN seen in tcpdump but no SYN-ACK, conntrack table full
4 TransportIs it listening? Does the handshake complete?ss -tlnp, nc -zv, tcpdump 'tcp port N', ss -tniNothing on the port, bound to 127.0.0.1, SYN retransmits, RST, retrans climbing, window 0
7 NamesDoes it resolve, correctly, for the user?dig, dig @resolver, getent hosts, resolvectl statusNXDOMAIN, SERVFAIL, timeout, a stale or geo-wrong answer, /etc/hosts override
6 TLSDoes the handshake complete and verify?curl -v, openssl s_client -connect host:443 -servername hostCertificate verify failed, stuck after Client Hello, protocol mismatch
7 HTTPWhat does the server say, and how fast?curl -sv -w, --resolve per backend, server logs502/503/504, slow time_starttransfer, one backend bad

The commands and how to read each output are in topic 18. The per-hop packet walk is in topic 2, the protocols in topic 6, TCP behaviour in topic 8.

Reference: what happens when you type https://www.facebook.com, and where it breaks

StepWhat happensWhere it failsSymptom you see
0 Host configDHCP earlier gave the host its address, mask, gateway and resolver.No DHCP answer169.254.x.x address, nothing works, ip addr shows it
1 BrowserParses the URL, checks HSTS and its own DNS and connection caches.Cached bad answer, forced HTTPS to a host with no TLSWorks in another browser or after a cache flush
2 DNSStub asks the recursive resolver; it walks root → .com → facebook.com authoritative; gets A/AAAA, often a short-TTL, geo-specific answer.Resolver unreachable; authoritative unreachable; zone wrong; geo mapping wrong"could not resolve host"; dig: timeout vs SERVFAIL vs NXDOMAIN; far-away IP returned, page slow
3 Route + ARPDestination is remote, so look up the default route and ARP for the gateway. Build the frame with the gateway's MAC.No default route; gateway not answering ARP; wrong VLAN"Network is unreachable" instantly; ip neigh FAILED; ping says Destination Host Unreachable
4 NAT and firewallHome or office edge rewrites source IP and port, creates a conntrack entry; stateful firewall remembers the outbound SYN.Translation table full; asymmetric return bypasses the deviceNew connections time out while existing ones keep working; conntrack -S insert_failed
5 TCPSYN to port 443 (IPv6 first if available), SYN-ACK, ACK. One RTT.SYN dropped by an ACL or firewall; nothing listening; a middlebox resettingHang then timeout (SYN retransmits at 1, 2, 4 s); immediate "connection refused" (RST); IPv6 broken but IPv4 fine
6 TLSClient Hello with SNI www.facebook.com and ALPN h2; server sends its certificate chain; keys agreed; one RTT in TLS 1.3.Wrong or expired cert, untrusted CA (corporate proxy), wrong client clock, MTU black hole eating the large Certificate message, protocol mismatchCertificate error page; "handshake failure"; curl -v stuck right after Client Hello
7 HTTPGET / with Host and cookies through an L4 load balancer to an L7 proxy to a backend.Backend down or slow; app error; rate limit; redirect loop502 or 504 (edge up, backend not); 503 (overloaded); 500 (app); 429; 301 loop
8 AssetsDozens of further connections to CDN hostnames for images and scripts.Different DNS name or edge failingPage loads, images and video missing

NAT, PAT and stateful firewalls (definitions)

Scenario 1: "A server is unreachable"

A user says: "I can't reach server S (10.9.5.5). Ping fails." Walk me through it.

Expected reasoning
  1. Scope. What does "unreachable" mean: ping, SSH, the web page, and the exact error? By name or by IP? Since when, and what changed? Then the three splits: one client or all clients, one server or all servers on that subnet, by IP or by name.
    • Name fails and IP works: DNS. dig, getent hosts, /etc/hosts. Done, different problem.
    • Many clients, one server: the server or its edge. One client, one server: the client, or something specific to that pair (ACL, route, host firewall).
    • One client, every server: the client's own stack. Start bottom-up on the client.
  2. Client, bottom-up. ip -br link (UP and LOWER_UP), ip addr (real address, not 169.254), ip route get 10.9.5.5 (which gateway and interface), ip neigh for that gateway (REACHABLE), ping gateway. If the gateway answers, the local link, VLAN and ARP are fine and the problem is at or beyond the gateway.
  3. Same subnet or remote? Same subnet: arping -I eth0 10.9.5.5. FAILED or no reply means a layer 2 fault: server down, wrong VLAN, cable, or two devices claiming the IP (two MACs reply). Remote: traceroute -n 10.9.5.5. The last responding hop is the furthest router that has both a route forward and a route back to you. The fault is at the next hop: no route to 10.9.0.0/16, an ACL, or no return route to the client's subnet.
  4. Server side (your shell, or a neighbour host on the same subnet). ip -br link carrier, ip addr right IP and mask, ip route default gateway, nft list ruleset. Then the decisive test: tcpdump -ni eth0 icmp and host <client> while the client pings.
    • Requests arrive, replies leave: the return path is broken. The server's route back is wrong, or a stateful firewall sees only one direction (asymmetric routing).
    • Requests arrive, no reply: the server's host firewall drops them, or rp_filter drops packets whose source is not routable back through the arriving interface.
    • Nothing arrives: the drop is upstream: an ACL, a missing route on an intermediate router, or the server's switch port in the wrong VLAN. Split again at the midpoint router.
  5. Is ping even the right test? ICMP may be filtered while the service is fine. nc -zv 10.9.5.5 22 or curl to the real port. "Server unreachable" and "ping blocked" are different incidents.
  6. Server down? Console or out-of-band: power, kernel panic, NIC driver, a reboot with a changed MAC while a neighbour holds a stale ARP entry. Check ip neigh on the gateway for 10.9.5.5.
  7. Report. "Confirmed: client stack and gateway fine, traceroute stops at R2. Ruled out: DNS, client firewall. Testing: route on R2 and return route on R3 next." Name the layer each test covered.

Scenario 2: "Users report latency and packet loss to a service across regions"

Users in region A say the service in region B is slow and drops. Possible reasons, and how do you narrow it?

Expected reasoning
  1. Baseline. What is normal RTT between A and B (light in fibre is roughly 1 ms per 100 km one way, so trans-Atlantic is 70 to 80 ms RTT) and what is it now? Is it all of region A or one site? Did synthetic monitoring see it, or only users? Since when, and what changed: a maintenance, a link flap, a BGP policy change, a new tunnel, a traffic shift?
  2. Measure both directions. mtr -rwzc 200 -n -T -P 443 from A toward B and from B toward A. TCP mode to the real port because ICMP is often deprioritised. Paths are often asymmetric, so one direction can be clean while the other is bad.
  3. Read mtr from the last hop backwards. Find the first hop whose loss and jitter persist to the destination; ignore hops whose loss vanishes downstream (ICMP rate limiting). A latency jump that persists means congestion at that hop or a path change (compare the hop list and AS path to yesterday's; a reroute through a longer path is latency with zero loss). Jitter (high StDev, high Wrst) with loss means queueing. Loss with flat latency means a bad link (CRC errors), a policer, or a black hole for some flows.
  4. Hop by hop at the suspect. Interface counters on both ends of that link: input errors and CRCs (dirty optic, bad cable, failing transceiver), output drops (queue full, congestion), utilisation graphs. On Linux: ethtool -S, ip -s link. On routers: the interface counters and queue drops. Errors climbing means replace the optic; drops with 95% utilisation means capacity or a traffic shift.
  5. Congestion vs loss vs reordering from the TCP side. ss -tni on the server for the affected flows: retrans climbing with cwnd small is loss; rtt growing before loss appears is a filling queue (congestion); dup ACKs and SACK blocks with no actual retransmits needed, or D-SACK, is reordering, which happens when parallel paths have different delays. Throughput ≈ MSS / (RTT × √loss), so even 0.5% loss at 150 ms kills a single flow (topic 8).
  6. MTU. Small transfers fine, big ones stall: ping -M do -s 1472 across the path. A new tunnel or MPLS segment with a smaller MTU and ICMP blocked is a black hole for large packets only, which users describe as "slow and flaky".
  7. ECMP: only some flows suffer. If some users are fine and others on the same path are not, suspect one member of an ECMP group or a LAG. Hashing on the 5-tuple sends each flow down one member, so a bad member hurts 1/N of flows. Test per flow by varying the source port: for p in $(seq 40000 40015); do mtr -rwc 50 -n -T -P 443 -L $p 10.9.5.5 | tail -1; done (or traceroute -T -p 443 --sport=$p). If two of sixteen show loss and the hop list for those shows a different member IP, you have found the bad link. Drain it (cost it out, shut the member), then fix the optic.
  8. Mitigate, then root-cause. Shift traffic off the bad path (local preference, cost-out, drain the link), confirm with the same mtr, keep the before and after. Then the postmortem: why did monitoring not catch it, and why did the link not auto-drain on errors.

Scenario 3: "The website is slow"

Users say the site is slow. The server team says CPU is fine. Where do you start?

Expected reasoning
  1. Define slow. Time to first byte, or full page load? Everyone, or one region or ISP? Every request, or intermittently? Which URL? Since when, and what changed?
  2. Reproduce with a timing breakdown from near the users and from near the server:
    curl -s -o /dev/null -w 'dns %{time_namelookup} tcp %{time_connect} tls %{time_appconnect} ttfb %{time_starttransfer} total %{time_total} code %{http_code}\n' https://www.example.com/
    Subtract adjacent values and the bucket names the suspect:
    • DNS large: slow or failing resolver, long CNAME chain, or geo-DNS sending users to a far edge.
    • TCP (connect minus namelookup) large: path RTT, or 1 s / 3 s steps, which are SYN retransmits, so a SYN is being lost.
    • TLS (appconnect minus connect) large: two RTTs (TLS 1.2 without resumption), OCSP fetch, or an overloaded TLS terminator.
    • TTFB (starttransfer minus appconnect) large: the server or something behind it (database, upstream API). This is the only bucket that is the server's.
    • Transfer (total minus starttransfer) large: the network during the download: loss, small windows, MTU.
  3. If transfer is slow, capture. tcpdump -w on the client and look for: retransmissions and dup ACKs (loss on the path; find it with mtr, scenario 2); window 0 from the client (the client or a proxy is not reading, so not a network problem); a small advertised window or no window scaling (buffer tuning); the first large response segment retransmitted over and over while the handshake and small requests are fine (PMTUD black hole; fix with MSS clamping or allow ICMP type 3 code 4). On the server, ss -tni shows retrans, cwnd and rtt per connection.
  4. If TTFB is slow, "CPU is fine" is not enough. Which server: the edge, the app tier, or the database? uptime and top for load, wa (disk) and si (softirq). Test each backend directly with curl --resolve host:443:backendIP: one slow backend behind a load balancer makes one request in N slow, which users report as "sometimes slow".
  5. CDN and DNS geo. Compare dig from the users' resolver with yours. If users' resolver maps them to a far edge (a resolver change, EDNS client subnet missing, a stale geo map), every connection pays a long RTT and the server is innocent. curl --resolve to the near edge versus the far edge quantifies it.
  6. Hand off with numbers. "DNS 20 ms, TCP 85 ms, TLS 90 ms, TTFB 1.9 s from Frankfurt against backend 3 only" sends the right team to the right place.

Scenario 4: "One host can't talk to one other host while neighbours are fine"

Host A (10.1.1.10) cannot reach host B (10.1.1.20). A reaches everything else. Other hosts reach B. What is going on?

Expected reasoning
  1. What the shape tells you. Pair-specific failure clears the core network, B's whole stack and A's whole stack. The fault is in how A and B see each other: addressing, ARP, a host firewall, a route, or a per-pair filter.
  2. Masks. ip addr on both. If A is 10.1.1.10/24 and B is 10.1.1.20/25, A treats B as on-link and ARPs directly, while B may treat A as remote and send its replies to the gateway (which may or may not forward them back onto the same segment). One direction works, the other depends on the gateway. A mask mismatch produces exactly "most things work, this pair does not".
  3. ARP. ip neigh show 10.1.1.20 on A. FAILED: B is not answering A on this link. REACHABLE with the wrong MAC: stale ARP after B's NIC or VM was replaced; ip neigh del and retry. arping -I eth0 10.1.1.20: two MACs replying means a duplicate IP; whichever answered last wins in each host's cache, so some hosts reach the real B and some reach the impostor. Check B for a PERMANENT entry pointing at A's old MAC.
  4. Host firewall on B. nft list ruleset or iptables -L -n -v on B, looking for a rule that names A's address. The most common real-world cause here is fail2ban or a similar tool that banned A after failed logins. Watch the DROP counter while A retries. Check A's OUTPUT chain too.
  5. VLAN. If A and B both reach remote hosts through their gateways but not each other, their switch ports may be in different VLANs that both carry a 10.1.1.0/24 address plan, or B's port was moved. ARP broadcasts do not cross VLANs. Check the switch: which VLAN each port is in, and on which port B's MAC was learned.
  6. ACL or security group. A per-host ACL, micro-segmentation policy or cloud security group that allows the subnet but not this pair. Prove it the same way as always: capture on B. A's packets arrive and no reply leaves means B (firewall, rp_filter). A's packets never arrive means something between: ACL on the switch, VLAN, or A sending them to the wrong place.
  7. Routes. ip route get 10.1.1.20 on A. It should say dev eth0 with no via. A stray /32 (a VPN client, a container network, an old static) sends that one destination somewhere else. Also ip rule for policy routing. Repeat on B toward A: the return path needs the same check.
  8. Decision table. ARP FAILED → L2 or VLAN. Wrong MAC → stale or duplicate. Packets arrive at B, no reply → B's firewall or rp_filter. Works one way only → mask mismatch or asymmetric route. Everything looks right but a /32 appears in ip route get → fix the route.

Two short ones

"DNS works but TCP doesn't."

dig returns an address; curl or nc -zv to it fails. DNS and the service are independent paths: the resolver is reachable, the server is not. First, is the address right? Compare with ip addr on the server; a stale record, split-horizon DNS or an /etc/hosts entry sends you to the wrong host. Then the failure type: refused (RST) means reachable but nothing listening on that address and port, so ss -tlnp on the server (bound to 127.0.0.1?); timeout means the SYN is dropped, so route, firewall or host down, which is the next item. Also check the address family: an AAAA record with broken IPv6 makes the client try v6 first; curl -4 versus -6 separates them.

"SYN goes out, no SYN-ACK comes back."

tcpdump on the client shows [S] retransmitted at 1, 2, 4, 8 s and nothing returning. Four possibilities, split by a capture on the server. (a) The SYN never reaches the server: no route, an ACL or firewall dropping silently, wrong VLAN. (b) It arrives and the server's host firewall drops it: tcpdump on the server shows the SYN, no SYN-ACK, DROP counter moves. (c) It arrives, the server answers, and the SYN-ACK is lost on the way back: wrong default route on the server, asymmetric path through a stateful firewall, rp_filter. (d) The listen backlog is full, so the kernel drops SYNs (nstat -az TcpExtListenDrops). An RST would be a different story: reachable, port closed.

Symptom vs cause: the October 2021 outage

Symptom: facebook.com, Instagram and WhatsApp stopped resolving; the world saw DNS failures. Cause: a routine backbone maintenance command, which an audit tool failed to stop, disconnected every data centre from the backbone; the authoritative DNS servers, designed to withdraw their BGP announcements when they cannot reach the data centres, did exactly that, so the DNS went unreachable too, and so did the internal tools and out-of-band access needed to fix it. Lesson: the first symptom you see is rarely the layer that broke; ask what the DNS servers depended on, and keep a path to your gear that does not depend on your gear. Details: engineering.fb.com, outage details. BGP mechanics are in topic 14.

Interview questions

1. Walk me through troubleshooting a server that is not reachable.Scope first: one client or all, one server or all, by name or by IP, since when, what changed. Then bottom-up from the client: link, address, route to the server (ip route get), ARP for the gateway, ping the gateway, traceroute to see where replies stop. Then from the server: is it up, right address and mask, default route, host firewall, and a capture to see whether the client's packets arrive and whether replies leave. Each result moves the fault to one half of the path. Follow-up they will ask: what if ping is blocked? Test the real port with nc or curl.
2. Users report latency and packet loss. What are possible reasons and how do you narrow it down?Reasons: congestion at a hop, a dirty link with CRC errors, a path change to a longer route, one bad ECMP or LAG member hurting some flows, an MTU black hole, or a problem on the server itself. Narrow it: baseline versus now, mtr in both directions with TCP probes, read loss from the last hop backwards to find the first hop where it persists, check error and drop counters on that link, and sweep source ports to catch a per-flow problem. Mitigate by draining the bad path, then fix it.
3. The website is slow. How do you approach it?Define slow and who sees it, then curl -w for the DNS, TCP, TLS, time-to-first-byte and transfer split from near the users. Each bucket points to a different cause: resolver or geo-DNS, path RTT or SYN loss, TLS round trips, the server and its backends, or loss and window size during transfer. Test individual backends with --resolve. Only the TTFB bucket is the server's, so "CPU is fine" answers one bucket of five.
4. Ping works but the service does not. What is going on?ICMP and TCP are filtered and routed independently. Likely: nothing listening on that port or bound only to localhost (RST, "refused"), a firewall allowing ICMP but dropping the port (timeout), an MTU black hole (handshake fine, data stalls), or a stateful device on an asymmetric path that passes ICMP but drops TCP state it never saw. ss -tlnp on the server and a capture settle it.
5. Give me the sixty-second version of what happens when I type https://www.facebook.com.Browser checks caches and HSTS. DNS: stub to recursive resolver, which walks root, .com and the authoritative servers and returns a short-TTL address near me. Host decides it is remote, finds the default route, ARPs for the gateway, sends the frame; each router rewrites MACs and decrements TTL, NAT rewrites my source at the edge. TCP three-way handshake to 443. TLS: Client Hello with SNI and ALPN, certificate chain, key agreement, one RTT in 1.3. HTTP GET through a load balancer to a backend, 200 with HTML, then dozens of connections for assets. Then the interviewer picks a step and goes deep: have each step's failure symptom ready from the table above.
6. What is NAT, and why does it break inbound connections?NAT rewrites addresses at a boundary, usually replacing many private source addresses with one public address and rewriting source ports so flows stay distinct (PAT). The translation table is built by outbound packets, so an inbound SYN with no matching entry has nowhere to go and is dropped unless a destination NAT rule maps it to an inside host. Side effects: idle entries expire, protocols that embed addresses need helpers, and the table can fill up.
7. Stateful versus stateless firewall, and why does asymmetric routing break the stateful one?A stateless ACL judges every packet alone and needs rules in both directions. A stateful firewall records the outbound SYN and automatically permits the matching return packets. If the reply comes back through a different device, that device has no state for the flow and drops it, so you get "SYN out, no SYN-ACK back" even though the server answered. The fix is symmetric routing or sharing state between the devices.
8. How do you prove the problem is the return path?Capture at the server: the client's packets arrive and the server's replies leave, yet the client never sees them. Then ip route get client on the server to see where replies go, and traceroute from the server toward the client to see where they stop. A forward traceroute alone can never show this, which is why you always test both directions.
9. One host cannot reach one other host, everybody else is fine. Your first three checks?Masks on both hosts (ip addr), the ARP entry for the peer on each side (ip neigh, arping for duplicates or a stale MAC), and the host firewall on the destination (iptables -L -n -v for a rule naming the source, often fail2ban). Then ip route get for a stray /32 and the switch for VLAN membership.
10. What does the October 2021 outage teach about troubleshooting?The visible symptom was DNS, but the cause was the backbone being disconnected by a maintenance command, with the DNS servers withdrawing their routes by design when they lost their data centres. So: do not stop at the first failing layer, ask what it depends on, and do not let your recovery path depend on the thing that broke, because the engineers lost internal tools and out-of-band access at the same time.

Traps

Scenario

You are handed this at the end of the interview: "A deploy went out an hour ago. Since then about one request in eight to the API fails with a timeout. Everything else is normal." What do you ask, and what do you test first?

Expected reasoning
  1. "What changed" is given: the deploy. Ask what it touched: application, config, number of backends, a kernel or network change. Ask whether the failing eighth is random or sticks to some users or some source addresses.
  2. One in eight is a ratio that smells like one bad member out of eight: one backend behind the load balancer, or one ECMP or LAG member. Random per request points to a backend; sticky per source address points to a hashed path.
  3. Test each backend directly with curl --resolve api.example.com:443:backendIP -w for all eight. One timing out or returning errors is the answer; check it with ss -tlnp, its logs and uptime. If all eight are fine from inside, sweep source ports through the load balancer and the path (scenario 2) to find a bad link.
  4. Mitigate as soon as you find it: drain that backend or link. Confirm the failure rate drops to zero with the same test. Then root-cause: why did the health check not catch it, and roll back or fix the deploy.
  5. Status along the way: "Confirmed: 1/8 timeouts since the deploy. Hypothesis: one backend. Testing each directly now; drain the bad one as soon as identified."

← 18 · Linux networking toolkit · all topics · 20 · The 45-minute coding interview script →