← all topics

18 · Linux networking toolkit Network

ip addr/link/route/neigh · ss · ping/traceroute/mtr · dig · curl -v · tcpdump · ethtool

Why it matters for NPE. NPEs live on Linux hosts. Interviewers ask which command answers which question and what its output proves.

Primer: every command answers one question

An NPE does not "run some commands". You ask a question about one layer, pick the command that answers it, read the one field that matters, and say out loud what that output proves and what it cannot prove. Interviewers grade that discipline more than the flags. This page is organised by question. Each tool gets a real-looking output with the fields you read annotated, so "what would you see" has a picture behind it.

The rule: a command proves something about the host you ran it on and the direction you ran it in. Everything else (the far end, the return path, the application) needs its own test. Say which it is every time.

Watch

How to Use the ip Command in Linux: A Beginner's GuideLearn Linux TV · 16:43

The first thing you type on a broken host. 2:38 basic usage, 3:51 one interface, 5:18 up/down, 6:43 add an address, 9:49 routing commands, 14:07 interface statistics. Skip 1:09 and 7:23 (sponsor). ip neigh is not covered; practise it from the section below.

Introduction to TCPDUMPDavid Mahler · 18:47

Terse and complete: -D -i -c -n -s0 -w/-r, host/port/net filters, -v/-X. The description lists every option used.

How to Use the ss Command (Linux Crash Course Series)Learn Linux TV · 6:56

2:05 basic usage, 3:31 UDP, 4:06 ss -s. Memorise ss -tlnp and ss -tni from this page; the video stops short.

How to Use the dig Command in Linux | DNS Lookup TutorialLearn Linux TV · 14:16

3:43 A records, 7:31 NS, 9:09 @server, 10:16 +short, 11:10 TTLs. Add +trace yourself.

Traceroute (tracert) Explained - Network TroubleshootingPowerCert Animated Videos · 9:23

The TTL-expiry mechanism in pictures. Nine minutes.

Traceroute explained // Featuring Elon Musk // Demo with Windows, Linux, macOSDavid Bombal · 22:35

With captures. 1:11 how it works, 5:21 TTL, 18:34 why Linux uses UDP and Windows uses ICMP, 19:56 Linux demo.

tcpdump - Traffic Capture & AnalysisHackerSploit · 23:20

Optional. Slower, more filter examples.

Basic cURL TutorialTraversy Media · 14:42

Only if you have never used curl. Reading beats video here.

Reading: everything curl, HTTP chapter (verbose output, -w, --resolve); Julia Evans' tcpdump zine; Cloudflare's reading mtr page.

Question → command → what it proves → what it can't

QuestionCommandWhat the output provesWhat it cannot prove
Is my interface up, with link, and what MTU?ip -br link, ip addr show eth0Admin state (UP), carrier (LOWER_UP), MTU, addresses and masks on this host.That the switch port is in the right VLAN, or that anything past the cable works.
Where would a packet to X leave from?ip route get XThe exact route, egress interface and source IP the kernel will use right now.That the next hop is alive, or that X has a route back.
Can I resolve my next hop at layer 2?ip neighREACHABLE/STALE: a MAC is known. FAILED: no ARP reply after retries.IP reachability. A REACHABLE entry can outlive a host that just died by up to 30 s.
Is anything listening? Is this connection healthy?ss -tlnp, ss -tniListening sockets with their process, connection states, queue depth, RTT, cwnd, retransmits.That a remote client can reach that port. Local firewall and upstream ACLs are invisible here.
Can I reach the host at layer 3?ping -cAn echo reply proves IP works in both directions right now. TTL in the reply gives hop count.That TCP to a port works. ICMP is often filtered or deprioritised separately. Ping loss is not TCP loss.
Where does the path stop or get lossy?traceroute -n, mtr -rwzc 100The forward hop sequence and per-hop latency as seen from here.The return path. A star is not loss. Per-hop loss counts only if it persists to the last hop.
Does the name resolve, to what, and who said so?dig, dig @server, dig +traceStatus, records, TTL, which server answered, whether the answer is authoritative.What the application will resolve: it uses /etc/hosts and nsswitch, dig does not.
Which stage of an HTTPS request is slow or failing?curl -v, curl -wSplit of DNS, TCP connect, TLS, time to first byte, transfer; cert verification; status code.Why the server took that long. That is the server's logs.
Did the packet actually leave or arrive?tcpdump -ni eth0Packets on the wire at this host: flags, seq, retransmits, who replied.Outbound packets dropped by the local firewall (captured after netfilter). Whether the far end got it. Capture there too.
Is the link dirty or is the host overrun?ip -s link, ethtool -S eth0Error and drop counters; errors/CRC = physical, dropped/missed/fifo = host could not drain the NIC.When it happened. Counters are cumulative: sample twice.
Can I open TCP to that port?nc -zv host portSucceeded: handshake done. Refused: RST, so reachable but nothing listening. Timeout: dropped somewhere.That the service behind the port is healthy.
Is the neighbour alive at L2? Is my IP duplicated?arping -I eth0 ip, arping -DARP replies and the MAC(s) that sent them, even when ICMP is filtered.Anything beyond the local link.
Is NAT or connection tracking eating it?conntrack -L, conntrack -STracked flows and their state; insert_failed/table full counters.Which rule dropped it: that is the ruleset.
Is the local firewall dropping it?nft list ruleset, iptables -S, iptables -L -n -vRules in order, default policy, and hit counters that move when you retry.Drops by an upstream ACL or security group.
Which resolver am I really using?cat /etc/resolv.conf, resolvectl status, getent hosts nameThe stub, the real upstream per link, and the answer an application gets via nsswitch.That the upstream resolver is healthy. Query it directly with dig.

Interfaces: ip link, ip addr

$ ip -br link
lo     UNKNOWN  00:00:00:00:00:00 <LOOPBACK,UP,LOWER_UP>
eth0   UP       52:54:00:1a:2b:3c <BROADCAST,MULTICAST,UP,LOWER_UP>
eth1   DOWN     52:54:00:4d:5e:6f <NO-CARRIER,BROADCAST,MULTICAST,UP>

$ ip addr show eth0
2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc mq state UP group default qlen 1000
    link/ether 52:54:00:1a:2b:3c brd ff:ff:ff:ff:ff:ff
    inet 10.1.1.10/24 brd 10.1.1.255 scope global eth0
    inet6 fe80::5054:ff:fe1a:2b3c/64 scope link

Routes: ip route, ip route get

$ ip route
default via 10.1.1.1 dev eth0 proto dhcp metric 100
10.1.1.0/24 dev eth0 proto kernel scope link src 10.1.1.10
10.9.0.0/16 via 10.1.1.254 dev eth0 proto static

$ ip route get 10.9.5.5
10.9.5.5 via 10.1.1.254 dev eth0 src 10.1.1.10 uid 1000
$ ip route get 10.1.1.20
10.1.1.20 dev eth0 src 10.1.1.10 uid 1000

ip route shows the table. ip route get asks the kernel the question you actually have: longest-prefix match done for you, including policy routing (ip rule). "via" means a gateway will be ARPed for; "dev eth0" with no via means the destination is treated as on-link. A stray more-specific route (a VPN client adding 10.1.1.20/32 dev tun0) shows up here and nowhere else you would think to look. RTNETLINK answers: Network is unreachable means no route at all, not a dead link.

Neighbours: ip neigh

$ ip neigh
10.1.1.1   dev eth0 lladdr 00:1c:73:aa:bb:cc REACHABLE
10.1.1.20  dev eth0 lladdr 52:54:00:99:88:77 STALE
10.1.1.77  dev eth0 FAILED
10.1.1.78  dev eth0 INCOMPLETE
StateMeaningWhat you do
REACHABLEMAC known and confirmed by traffic within the last ~30 s.Trust it; L2 to that IP works.
STALEMAC known, not recently confirmed. Still used; the kernel sends a unicast probe on next use.Normal. Not a fault.
DELAY / PROBETransitional while re-verifying.Ignore.
INCOMPLETEARP request sent, no reply yet.Watch it. It becomes FAILED or REACHABLE.
FAILEDNo ARP reply after retries.The host did not answer on this link: down, wrong VLAN, wrong subnet, or a different L2 segment.
PERMANENTStatic entry added by hand.Suspect it after a NIC swap: the MAC changed and this did not.

Clear a bad entry with ip neigh del 10.1.1.20 dev eth0 or ip neigh flush dev eth0. Prove the L2 path independently with arping -I eth0 -c 3 10.1.1.20: two different MACs replying means a duplicate IP.

Sockets: ss

$ ss -tlnp
State   Recv-Q  Send-Q  Local Address:Port  Peer Address:Port  Process
LISTEN  0       4096    0.0.0.0:22          0.0.0.0:*          users:(("sshd",pid=812,fd=3))
LISTEN  0       511     127.0.0.1:8080      0.0.0.0:*          users:(("gunicorn",pid=2210,fd=5))
LISTEN  0       4096    [::]:443            [::]:*             users:(("nginx",pid=1407,fd=6))

$ ss -tan state syn-sent
Recv-Q Send-Q  Local Address:Port   Peer Address:Port
0      1       10.1.1.10:40122      10.9.5.5:5432

$ ss -tni dst 203.0.113.5
ESTAB 0 0 10.1.1.10:51234 203.0.113.5:443
     cubic wscale:7,7 rto:204 rtt:3.1/0.5 mss:1448 pmtu:1500 cwnd:10 bytes_sent:4210
     bytes_acked:4210 bytes_received:18220 segs_out:31 segs_in:29 retrans:0/2
     lastsnd:120 lastrcv:118 delivery_rate 29.9Mbps minrtt:2.8

Reachability: ping

$ ping -c 4 -i 0.2 10.9.5.5
PING 10.9.5.5 (10.9.5.5) 56(84) bytes of data.
64 bytes from 10.9.5.5: icmp_seq=1 ttl=61 time=1.84 ms
64 bytes from 10.9.5.5: icmp_seq=2 ttl=61 time=1.79 ms
64 bytes from 10.9.5.5: icmp_seq=4 ttl=61 time=12.3 ms

--- 10.9.5.5 ping statistics ---
4 packets transmitted, 3 received, 25% packet loss, time 612ms
rtt min/avg/max/mdev = 1.790/5.310/12.300/4.943 ms

$ ping -M do -s 1472 -c 1 10.9.5.5      # 1472 + 8 ICMP + 20 IP = 1500 bytes
$ ping -M do -s 1473 -c 1 10.9.5.5
ping: local error: message too long, mtu=1500
$ ping -M do -s 1400 -c 1 10.9.5.5      # across a tunnel
From 10.1.1.254 icmp_seq=1 Frag needed and DF set (mtu = 1400)

Path: traceroute and mtr

Both send probes with TTL 1, 2, 3 and so on. Each router that decrements TTL to zero drops the probe and returns ICMP Time Exceeded (type 11) from its own address. The destination answers differently: ICMP Port Unreachable (UDP mode), Echo Reply (ICMP mode) or SYN-ACK/RST (TCP mode). Linux defaults to UDP to ports 33434 and up, Windows uses ICMP, -I forces ICMP, -T -p 443 uses TCP SYNs so the probes follow whatever firewalls allow for the real traffic.

$ traceroute -n 203.0.113.5
traceroute to 203.0.113.5 (203.0.113.5), 30 hops max, 60 byte packets
 1  10.1.1.1        0.412 ms  0.388 ms  0.371 ms
 2  10.0.0.9        1.102 ms  1.080 ms  1.055 ms
 3  * * *
 4  198.51.100.14   8.913 ms  8.870 ms  8.851 ms
 5  203.0.113.5     9.204 ms  9.180 ms  9.150 ms

$ mtr -rwzc 100 -n -T -P 443 203.0.113.5
HOST: client01                 Loss%   Snt   Last   Avg  Best  Wrst StDev
  1. AS???    10.1.1.1          0.0%   100    0.4   0.4   0.3   0.9   0.1
  2. AS64500  10.0.0.9          0.0%   100    1.1   1.1   1.0   2.3   0.2
  3. AS???    ???             100.0%   100    0.0   0.0   0.0   0.0   0.0
  4. AS64501  198.51.100.14    40.0%   100    8.9   9.0   8.8  11.2   0.4
  5. AS64501  198.51.100.22     3.0%   100    9.1  14.8   9.0  88.0  15.1
  6. AS64502  203.0.113.5       3.0%   100    9.2  15.1   9.1  90.2  15.4

Names: dig

$ dig www.example.com

;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 31337
;; flags: qr rd ra; QUERY: 1, ANSWER: 2, AUTHORITY: 0, ADDITIONAL: 1

;; QUESTION SECTION:
;www.example.com.            IN  A

;; ANSWER SECTION:
www.example.com.     300  IN  CNAME  edge.example.net.
edge.example.net.     30  IN  A      203.0.113.5

;; Query time: 23 msec
;; SERVER: 10.1.1.53#53(10.1.1.53) (UDP)

$ dig +short www.example.com @8.8.8.8
edge.example.net.
203.0.113.5
$ dig +trace www.example.com | grep -E '^(www|example|com)'
FieldRead it as
status: NOERRORThe name exists. With ANSWER: 0 it exists but has no record of that type (asked AAAA, only A exists).
status: NXDOMAINThe authoritative server says the name does not exist. Typo, missing record, wrong search domain, or a broken zone.
status: SERVFAILThe resolver could not get an answer: authoritative servers unreachable, DNSSEC validation failure, lame delegation. The resolver is up; its upstream is the problem.
status: REFUSEDPolicy: that server will not answer you (wrong ACL, asking an authoritative server for recursion).
;; connection timed out; no servers could be reachedYou never reached the resolver. This is a network or resolver-down problem, not a DNS data problem.
flags: aaAuthoritative answer. Only present when you ask the authoritative server directly. ra = recursion available, rd = recursion desired.
ANSWER vs AUTHORITYANSWER holds the records you asked for. AUTHORITY holds the NS records (or SOA on a negative answer) telling you who owns the zone.
TTL (300, 30)Seconds the answer may be cached. A low TTL on the A record is a load balancer or CDN moving traffic. A stale cache during a change shows as "my dig and yours differ": compare with @server.
+traceWalks root → .com → example.com authoritative yourself, bypassing every cache. Use it when resolvers disagree.

dig -x 203.0.113.5 does the PTR lookup. Remember that dig speaks only DNS: getent hosts www.example.com goes through nsswitch (/etc/hosts first, then DNS) the way an application does. Resolution chain detail is in topic 6.

HTTP(S): curl -v and -w

$ curl -sv -o /dev/null https://www.example.com/ 2>&1 | grep -E '^(\*|>|<) ' | head -20
* Host www.example.com:443 was resolved.
* IPv4: 203.0.113.5
*   Trying 203.0.113.5:443...
* Connected to www.example.com (203.0.113.5) port 443
* ALPN: curl offers h2,http/1.1
* TLSv1.3 (OUT), TLS handshake, Client hello (1):
* TLSv1.3 (IN), TLS handshake, Server hello (2):
* TLSv1.3 (IN), TLS handshake, Certificate (11):
* SSL connection using TLSv1.3 / TLS_AES_256_GCM_SHA384
* Server certificate:
*  subject: CN=www.example.com
*  SSL certificate verify ok.
* using HTTP/2
> GET / HTTP/2
> Host: www.example.com
< HTTP/2 200
< content-type: text/html

$ curl -s -o /dev/null -w 'dns %{time_namelookup} tcp %{time_connect} tls %{time_appconnect} ttfb %{time_starttransfer} total %{time_total} code %{http_code}\n' https://www.example.com/
dns 0.021 tcp 0.024 tls 0.061 ttfb 0.412 total 0.418 code 200

The wire: tcpdump

$ sudo tcpdump -ni eth0 -c 5 'host 10.9.5.5 and tcp port 5432'
listening on eth0, link-type EN10MB (Ethernet), snapshot length 262144 bytes
12:00:01.000112 IP 10.1.1.10.40122 > 10.9.5.5.5432: Flags [S], seq 1998456112, win 64240, options [mss 1460,sackOK,TS val 1 ecr 0,nop,wscale 7], length 0
12:00:02.021330 IP 10.1.1.10.40122 > 10.9.5.5.5432: Flags [S], seq 1998456112, win 64240, options [mss 1460,sackOK,TS val 1021 ecr 0,nop,wscale 7], length 0
12:00:04.069420 IP 10.1.1.10.40122 > 10.9.5.5.5432: Flags [S], seq 1998456112, win 64240, options [mss 1460,sackOK,TS val 3069 ecr 0,nop,wscale 7], length 0

$ sudo tcpdump -ni any -s0 -w /tmp/case.pcap 'host 10.9.5.5 and not port 22'
$ sudo tcpdump -nr /tmp/case.pcap 'tcp[tcpflags] & (tcp-syn|tcp-rst) != 0'
FlagMeaningFlagMeaning
[S]SYN[S.]SYN-ACK (the dot is ACK)
[.]ACK only[P.]PSH-ACK, data being delivered
[F.]FIN-ACK, polite close[R] / [R.]RST: closed port, firewall reject, or abort

Counters: ip -s link, ethtool -S

$ ip -s link show eth0
2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc mq state UP mode DEFAULT group default qlen 1000
    RX:  bytes    packets  errors  dropped  missed  mcast
    9812345678   8123456    1532    12044       0   4410
    TX:  bytes    packets  errors  dropped  carrier collsns
    4123456789   5123456       0        0       0       0

$ ethtool -S eth0 | grep -Ei 'err|drop|miss|crc|over|fifo' | grep -v ': 0$'
     rx_crc_errors: 1532
     rx_missed_errors: 12044
     rx_fifo_errors: 12044
$ ethtool eth0 | grep -E 'Speed|Duplex|Link detected'
	Speed: 10000Mb/s
	Duplex: Full
	Link detected: yes

One-liners you should be able to say

ToolLineReading
ncnc -zv 10.9.5.5 5432"succeeded" = handshake done. "Connection refused" = RST came back: reachable, nothing listening. Hang then timeout = SYN dropped or no route. UDP (-u) can only prove "closed" when an ICMP port unreachable comes back.
arpingarping -I eth0 -c 3 10.1.1.20, arping -D -I eth0 10.1.1.10L2 reachability even when ICMP is filtered. -D is duplicate address detection: any reply means someone else has your IP. Two MACs replying to the first form means a duplicate on the segment.
conntrackconntrack -L | grep 10.9.5.5, conntrack -SThe kernel's flow table behind NAT and stateful rules: state per flow, and insert_failed/drop counters. "nf_conntrack: table full, dropping packet" in dmesg explains new connections failing while old ones work.
nft / iptablesnft list ruleset, iptables -S, iptables -L -n -vRules in order and the chain policy (ACCEPT or DROP). -v shows packet counters per rule: retry the failing connection and watch which counter moves. Also check sysctl net.ipv4.conf.all.rp_filter if packets arrive with "wrong" source routes.
/etc/resolv.confcat /etc/resolv.confnameserver 127.0.0.53 is the systemd-resolved stub; the real upstream is in resolvectl status. search adds suffixes to short names. options timeout:2 attempts:2 bounds how long a dead resolver can hurt you.
/etc/hostsgrep name /etc/hosts, getent hosts nameWins over DNS when /etc/nsswitch.conf says hosts: files dns. A stale entry here makes dig and the application disagree.
systemd-resolvesystemd-resolve --status (older) or resolvectl status, resolvectl query name, resolvectl flush-cachesPer-link DNS servers and search domains, which link a query will use, and a cache you can flush to rule out stale answers.

Linux host health in one pass

NPE interviews are not the PE Linux round, but "what causes high CPU" and "how do you find a Python memory leak" have been asked. Have one command and one sentence per resource.

CPU

$ uptime
 14:02:11 up 41 days,  3:12,  2 users,  load average: 9.84, 6.12, 2.40
$ nproc
8
$ ps -eo pcpu,pmem,pid,etime,cmd --sort=-pcpu | head -4
%CPU %MEM   PID     ELAPSED CMD
98.2  1.1  2210  3-04:11:02 /usr/bin/python3 collector.py
 3.0  0.4  1407    41-03:12 nginx: worker process
 0.4  0.1   812    41-03:12 sshd: /usr/sbin/sshd

Memory

$ free -m
               total   used   free  shared  buff/cache  available
Mem:           32000  21500    900     120        9600      10100
Swap:           2047    800   1247
$ dmesg -T | grep -i 'killed process'
[Tue ...] Out of memory: Killed process 2210 (python3) total-vm:9123456kB, anon-rss:8234560kB

Finding a Python memory leak

$ while true; do ps -o rss= -p 2210; sleep 60; done      # RSS in kB, does it plateau or climb forever?

import tracemalloc
tracemalloc.start(25)
s1 = tracemalloc.take_snapshot()
# ... run the workload for a while ...
s2 = tracemalloc.take_snapshot()
for stat in s2.compare_to(s1, 'lineno')[:5]:
    print(stat)
# collector.py:88: size=412 MiB (+398 MiB), count=3100000 (+3000000), average=139 B

Disk, logs, file descriptors

$ df -h /                 # Use% 100%: "No space left on device"
$ df -i /                 # inodes can fill while bytes are free
$ du -xsh /var/log/* 2>/dev/null | sort -rh | head
$ lsof +L1                # deleted files still held open: space not freed until the process closes them

$ journalctl -u nginx -S -1h -p err --no-pager
$ journalctl -k -S -30min   # kernel: link up/down, OOM, conntrack full
$ dmesg -T | tail -50

$ ls /proc/2210/fd | wc -l ; grep 'open files' /proc/2210/limits
$ lsof -p 2210 | awk '{print $5}' | sort | uniq -c | sort -rn | head
$ ulimit -n                 # this shell only; services get LimitNOFILE= from systemd

Interview questions

1. What does a star in traceroute mean?No reply came back for that probe within the timeout. The usual reason is that the router at that hop does not generate ICMP Time Exceeded, rate-limits it, or a filter drops it; MPLS cores often hide hops this way. If later hops answer, traffic is flowing through the starred hop and it is not a problem. Only stars that continue to the end suggest a real break, and even then the destination may just be filtering the probe type, so retry with -I or -T -p 443.
2. ss shows a socket in SYN-SENT. What does that tell you?The host sent a SYN and received neither a SYN-ACK nor a RST. So the far end is silent to us: no route, a silent firewall drop in either direction, the server down, or its reply lost on the return path. It rules out "port closed", because a closed port would answer RST and the socket would be gone. Next step: tcpdump on the client to see the SYN retransmits leave, then a capture on the server to see whether they arrive.
3. How would you confirm a packet actually left the host?tcpdump -ni eth0 filtered on the destination, while generating the traffic. If the packet appears on the egress interface, it left. Two caveats: tcpdump sees outbound packets after netfilter, so a local OUTPUT drop means it never appears, and seeing it leave proves nothing about arrival. Pair it with ip route get to confirm it went out the interface you expected, and -e to check the destination MAC is the right gateway.
4. What is the difference between UP and LOWER_UP on an interface?UP is the administrative state: the interface is enabled in software. LOWER_UP is the carrier: the NIC detects a live link on the cable. UP without LOWER_UP (shown as NO-CARRIER) is a physical problem or a down switch port; LOWER_UP without an address or route is a configuration problem.
5. What do Recv-Q and Send-Q mean?On a LISTEN socket, Recv-Q is the number of completed connections waiting for accept() and Send-Q is the backlog limit; Recv-Q near Send-Q means the app is not accepting and SYNs are being dropped. On an established socket, Recv-Q is bytes received that the application has not read (slow app) and Send-Q is bytes sent but not yet acknowledged (slow network or peer).
6. dig returns NXDOMAIN, SERVFAIL, or times out. What does each mean?NXDOMAIN: the authoritative server says the name does not exist, so look at the zone or the spelling. SERVFAIL: the resolver tried and failed, usually because the authoritative servers are unreachable or DNSSEC validation failed. Timeout: you never reached the resolver at all, which is a network or resolver-down problem. Follow-up: ask the same question of a different resolver with @8.8.8.8 and of the authoritative server with +trace to see who disagrees.
7. How do you find the path MTU with ping?Send pings with the DF bit set (-M do) and a payload size, starting at 1472 (1500 minus 28 bytes of headers). If a link on the path is smaller, the router returns ICMP Fragmentation Needed with its MTU, and you step down. If large probes silently disappear while small ones work, something is dropping that ICMP, which is the PMTUD black hole from topic 8.
8. mtr shows 40% loss at hop 4 and 0% at the destination. Is there a problem?No. Loss at an intermediate hop that does not continue to later hops is that router rate-limiting its own ICMP replies; transit traffic through it is fine. Real loss starts at some hop and persists to the final hop, because every later probe has to cross the bad link. Read the table from the bottom up.
9. A page takes two seconds. How do you tell DNS from network from server?curl -w with time_namelookup, time_connect, time_appconnect, time_starttransfer and time_total, then subtract adjacent values. DNS is the first number. Connect minus namelookup is one RTT, so a value of one or three seconds is a lost SYN. TLS is appconnect minus connect, one RTT in TLS 1.3. Starttransfer minus appconnect is the server's think time. The remainder is transfer, where loss and window size show up. Each bucket points to a different team.
10. Load average is 12 on an 8-core server. What does that mean, and what do you type next?On average 12 tasks were runnable or stuck in uninterruptible I/O over the last minute, so the box is oversubscribed. Compare the 1, 5 and 15 minute numbers to see if it is rising. Then top for the us/sy/wa/si split (user code vs kernel vs disk vs softirq) and ps --sort=-pcpu or pidstat to name the process. High wa with low us means it is a disk problem wearing a CPU costume.

Traps

Scenario

A health check against a web server started failing ten minutes ago. You have a shell on the server and nothing else. Which commands, in which order, and what does each one rule out?

Expected reasoning
  1. ss -tlnp | grep ':443': is anything listening, and on which address? Nothing listening means the service died: systemctl status, journalctl -u nginx -S -15min, dmesg -T | grep -i killed for the OOM killer. Listening on 127.0.0.1 only means a config change broke the bind address.
  2. curl -sv -o /dev/null https://localhost/ --resolve www.example.com:443:127.0.0.1: does the service answer locally? A 200 here moves the fault off the application and onto the path in. A 502 or a hang keeps it on the box (backend, disk full, CPU).
  3. ip -br link, ip addr, ip route get <health-checker IP>: link up with carrier, the right address and mask, a route back to the checker. No route back is a complete explanation on its own.
  4. tcpdump -ni eth0 'host <checker> and port 443' while the check runs. SYNs arrive and SYN-ACKs leave: the server is fine, the return path or the checker is the problem. SYNs arrive, no SYN-ACK: local firewall (nft list ruleset, look for a new DROP, check counters) or the listener is gone. Nothing arrives: upstream, so escalate with "the box is healthy and sees no traffic".
  5. uptime, free -m, df -h take ten seconds and catch the boring causes: a disk that filled ten minutes ago, a load spike, an OOM.
  6. Report in one line: what is confirmed, what is ruled out, what you are testing next. The structure is the answer; the commands are the evidence.

← 17 · Classics: binary search, recursion, trees and BST · all topics · 19 · Troubleshooting walkthroughs →