18 · Linux networking toolkit Network
ip addr/link/route/neigh · ss · ping/traceroute/mtr · dig · curl -v · tcpdump · ethtool
Why it matters for NPE. NPEs live on Linux hosts. Interviewers ask which command answers which question and what its output proves.
Primer: every command answers one question
An NPE does not "run some commands". You ask a question about one layer, pick the command that answers it, read the one field that matters, and say out loud what that output proves and what it cannot prove. Interviewers grade that discipline more than the flags. This page is organised by question. Each tool gets a real-looking output with the fields you read annotated, so "what would you see" has a picture behind it.
Watch
The first thing you type on a broken host. 2:38 basic usage, 3:51 one interface, 5:18 up/down, 6:43 add an address, 9:49 routing commands, 14:07 interface statistics. Skip 1:09 and 7:23 (sponsor). ip neigh is not covered; practise it from the section below.
Terse and complete: -D -i -c -n -s0 -w/-r, host/port/net filters, -v/-X. The description lists every option used.
2:05 basic usage, 3:31 UDP, 4:06 ss -s. Memorise ss -tlnp and ss -tni from this page; the video stops short.
3:43 A records, 7:31 NS, 9:09 @server, 10:16 +short, 11:10 TTLs. Add +trace yourself.
The TTL-expiry mechanism in pictures. Nine minutes.
With captures. 1:11 how it works, 5:21 TTL, 18:34 why Linux uses UDP and Windows uses ICMP, 19:56 Linux demo.
Only if you have never used curl. Reading beats video here.
Reading: everything curl, HTTP chapter (verbose output, -w, --resolve); Julia Evans' tcpdump zine; Cloudflare's reading mtr page.
Question → command → what it proves → what it can't
| Question | Command | What the output proves | What it cannot prove |
|---|---|---|---|
| Is my interface up, with link, and what MTU? | ip -br link, ip addr show eth0 | Admin state (UP), carrier (LOWER_UP), MTU, addresses and masks on this host. | That the switch port is in the right VLAN, or that anything past the cable works. |
| Where would a packet to X leave from? | ip route get X | The exact route, egress interface and source IP the kernel will use right now. | That the next hop is alive, or that X has a route back. |
| Can I resolve my next hop at layer 2? | ip neigh | REACHABLE/STALE: a MAC is known. FAILED: no ARP reply after retries. | IP reachability. A REACHABLE entry can outlive a host that just died by up to 30 s. |
| Is anything listening? Is this connection healthy? | ss -tlnp, ss -tni | Listening sockets with their process, connection states, queue depth, RTT, cwnd, retransmits. | That a remote client can reach that port. Local firewall and upstream ACLs are invisible here. |
| Can I reach the host at layer 3? | ping -c | An echo reply proves IP works in both directions right now. TTL in the reply gives hop count. | That TCP to a port works. ICMP is often filtered or deprioritised separately. Ping loss is not TCP loss. |
| Where does the path stop or get lossy? | traceroute -n, mtr -rwzc 100 | The forward hop sequence and per-hop latency as seen from here. | The return path. A star is not loss. Per-hop loss counts only if it persists to the last hop. |
| Does the name resolve, to what, and who said so? | dig, dig @server, dig +trace | Status, records, TTL, which server answered, whether the answer is authoritative. | What the application will resolve: it uses /etc/hosts and nsswitch, dig does not. |
| Which stage of an HTTPS request is slow or failing? | curl -v, curl -w | Split of DNS, TCP connect, TLS, time to first byte, transfer; cert verification; status code. | Why the server took that long. That is the server's logs. |
| Did the packet actually leave or arrive? | tcpdump -ni eth0 | Packets on the wire at this host: flags, seq, retransmits, who replied. | Outbound packets dropped by the local firewall (captured after netfilter). Whether the far end got it. Capture there too. |
| Is the link dirty or is the host overrun? | ip -s link, ethtool -S eth0 | Error and drop counters; errors/CRC = physical, dropped/missed/fifo = host could not drain the NIC. | When it happened. Counters are cumulative: sample twice. |
| Can I open TCP to that port? | nc -zv host port | Succeeded: handshake done. Refused: RST, so reachable but nothing listening. Timeout: dropped somewhere. | That the service behind the port is healthy. |
| Is the neighbour alive at L2? Is my IP duplicated? | arping -I eth0 ip, arping -D | ARP replies and the MAC(s) that sent them, even when ICMP is filtered. | Anything beyond the local link. |
| Is NAT or connection tracking eating it? | conntrack -L, conntrack -S | Tracked flows and their state; insert_failed/table full counters. | Which rule dropped it: that is the ruleset. |
| Is the local firewall dropping it? | nft list ruleset, iptables -S, iptables -L -n -v | Rules in order, default policy, and hit counters that move when you retry. | Drops by an upstream ACL or security group. |
| Which resolver am I really using? | cat /etc/resolv.conf, resolvectl status, getent hosts name | The stub, the real upstream per link, and the answer an application gets via nsswitch. | That the upstream resolver is healthy. Query it directly with dig. |
Interfaces: ip link, ip addr
$ ip -br link
lo UNKNOWN 00:00:00:00:00:00 <LOOPBACK,UP,LOWER_UP>
eth0 UP 52:54:00:1a:2b:3c <BROADCAST,MULTICAST,UP,LOWER_UP>
eth1 DOWN 52:54:00:4d:5e:6f <NO-CARRIER,BROADCAST,MULTICAST,UP>
$ ip addr show eth0
2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc mq state UP group default qlen 1000
link/ether 52:54:00:1a:2b:3c brd ff:ff:ff:ff:ff:ff
inet 10.1.1.10/24 brd 10.1.1.255 scope global eth0
inet6 fe80::5054:ff:fe1a:2b3c/64 scope link
- UP in the flags is administrative: someone ran
ip link set up. LOWER_UP is the carrier: the NIC sees light or link pulses.eth1is admin up with NO-CARRIER: cable, optic, or the far switch port is down.state DOWNalone would mean admin down. - mtu 1500: compare to the path. A 9000 here and 1500 on the far side gives a black hole for big packets (topic 8).
- Check the mask: /24 vs /25 decides whether the host ARPs for the peer or sends to the gateway (topic 2).
ethtool eth0adds speed, duplex and "Link detected". A 10G port negotiated to 1G or half duplex is a physical finding.
Routes: ip route, ip route get
$ ip route
default via 10.1.1.1 dev eth0 proto dhcp metric 100
10.1.1.0/24 dev eth0 proto kernel scope link src 10.1.1.10
10.9.0.0/16 via 10.1.1.254 dev eth0 proto static
$ ip route get 10.9.5.5
10.9.5.5 via 10.1.1.254 dev eth0 src 10.1.1.10 uid 1000
$ ip route get 10.1.1.20
10.1.1.20 dev eth0 src 10.1.1.10 uid 1000
ip route shows the table. ip route get asks the kernel the question you actually have: longest-prefix match done for you, including policy routing (ip rule). "via" means a gateway will be ARPed for; "dev eth0" with no via means the destination is treated as on-link. A stray more-specific route (a VPN client adding 10.1.1.20/32 dev tun0) shows up here and nowhere else you would think to look. RTNETLINK answers: Network is unreachable means no route at all, not a dead link.
Neighbours: ip neigh
$ ip neigh
10.1.1.1 dev eth0 lladdr 00:1c:73:aa:bb:cc REACHABLE
10.1.1.20 dev eth0 lladdr 52:54:00:99:88:77 STALE
10.1.1.77 dev eth0 FAILED
10.1.1.78 dev eth0 INCOMPLETE
| State | Meaning | What you do |
|---|---|---|
| REACHABLE | MAC known and confirmed by traffic within the last ~30 s. | Trust it; L2 to that IP works. |
| STALE | MAC known, not recently confirmed. Still used; the kernel sends a unicast probe on next use. | Normal. Not a fault. |
| DELAY / PROBE | Transitional while re-verifying. | Ignore. |
| INCOMPLETE | ARP request sent, no reply yet. | Watch it. It becomes FAILED or REACHABLE. |
| FAILED | No ARP reply after retries. | The host did not answer on this link: down, wrong VLAN, wrong subnet, or a different L2 segment. |
| PERMANENT | Static entry added by hand. | Suspect it after a NIC swap: the MAC changed and this did not. |
Clear a bad entry with ip neigh del 10.1.1.20 dev eth0 or ip neigh flush dev eth0. Prove the L2 path independently with arping -I eth0 -c 3 10.1.1.20: two different MACs replying means a duplicate IP.
Sockets: ss
$ ss -tlnp
State Recv-Q Send-Q Local Address:Port Peer Address:Port Process
LISTEN 0 4096 0.0.0.0:22 0.0.0.0:* users:(("sshd",pid=812,fd=3))
LISTEN 0 511 127.0.0.1:8080 0.0.0.0:* users:(("gunicorn",pid=2210,fd=5))
LISTEN 0 4096 [::]:443 [::]:* users:(("nginx",pid=1407,fd=6))
$ ss -tan state syn-sent
Recv-Q Send-Q Local Address:Port Peer Address:Port
0 1 10.1.1.10:40122 10.9.5.5:5432
$ ss -tni dst 203.0.113.5
ESTAB 0 0 10.1.1.10:51234 203.0.113.5:443
cubic wscale:7,7 rto:204 rtt:3.1/0.5 mss:1448 pmtu:1500 cwnd:10 bytes_sent:4210
bytes_acked:4210 bytes_received:18220 segs_out:31 segs_in:29 retrans:0/2
lastsnd:120 lastrcv:118 delivery_rate 29.9Mbps minrtt:2.8
- Flags:
-tTCP,-uUDP,-llistening,-nnumeric (always),-pprocess,-aall,-iTCP internals,-ssummary.state established,state time-waitanddst/dportfilters narrow it. - gunicorn on 127.0.0.1:8080 is the classic "works on the box, refused from outside". The bind address matters.
- Recv-Q / Send-Q on LISTEN: Recv-Q is the number of completed connections waiting for
accept(); Send-Q is the backlog limit. Recv-Q near Send-Q means the application is not accepting fast enough and new SYNs get dropped (nstat -az TcpExtListenOverflows). - Recv-Q / Send-Q on ESTAB: Recv-Q is bytes received and not yet read by the application (the app is slow). Send-Q is bytes sent and not yet acknowledged (the network or the peer is slow).
- States: LISTEN waiting; SYN-SENT means we sent a SYN and got nothing back yet, neither SYN-ACK nor RST, so the far side is silent or unreachable; SYN-RECV on a server is the half-open backlog; ESTAB; FIN-WAIT, CLOSE-WAIT (our app has not called close), TIME-WAIT (normal on a busy client, 60 s). States are in topic 8.
-ifields:rtt:3.1/0.5is smoothed RTT and variance in ms;rto:204the retransmission timer;cwnd:10segments in flight allowed (stuck at a few = loss);retrans:0/2is current/total retransmits;pmtu:1500the path MTU the stack believes;delivery_ratethe goodput estimate. A connection withretransclimbing andcwndsmall is losing packets; one with large Recv-Q and zero retrans has a slow application.
Reachability: ping
$ ping -c 4 -i 0.2 10.9.5.5
PING 10.9.5.5 (10.9.5.5) 56(84) bytes of data.
64 bytes from 10.9.5.5: icmp_seq=1 ttl=61 time=1.84 ms
64 bytes from 10.9.5.5: icmp_seq=2 ttl=61 time=1.79 ms
64 bytes from 10.9.5.5: icmp_seq=4 ttl=61 time=12.3 ms
--- 10.9.5.5 ping statistics ---
4 packets transmitted, 3 received, 25% packet loss, time 612ms
rtt min/avg/max/mdev = 1.790/5.310/12.300/4.943 ms
$ ping -M do -s 1472 -c 1 10.9.5.5 # 1472 + 8 ICMP + 20 IP = 1500 bytes
$ ping -M do -s 1473 -c 1 10.9.5.5
ping: local error: message too long, mtu=1500
$ ping -M do -s 1400 -c 1 10.9.5.5 # across a tunnel
From 10.1.1.254 icmp_seq=1 Frag needed and DF set (mtu = 1400)
-ccount,-iinterval (below 0.2 s needs root),-spayload size,-M dosets the DF bit so routers must drop instead of fragment,-Wtimeout,-Isource interface or address.- ttl=61 from a Linux host (starts at 64) means 3 routers in the return path. Windows starts at 128, most network gear at 255.
- Four packets prove nothing about loss. 25% here is one lost packet. Use
-c 100or more, and remember ICMP is often rate limited or deprioritised by routers, so ping loss can exist with zero TCP loss. Destination Host Unreachablefrom your own IP means ARP failed on the local link.Network is unreachablemeans no route. Silence means dropped somewhere, or the far side filters ICMP.- MTU probing: step the size until "Frag needed" appears; the
mtu =value is the small link. If large probes just vanish with no ICMP, a firewall is blocking the ICMP and you have found a PMTUD black hole.
Path: traceroute and mtr
Both send probes with TTL 1, 2, 3 and so on. Each router that decrements TTL to zero drops the probe and returns ICMP Time Exceeded (type 11) from its own address. The destination answers differently: ICMP Port Unreachable (UDP mode), Echo Reply (ICMP mode) or SYN-ACK/RST (TCP mode). Linux defaults to UDP to ports 33434 and up, Windows uses ICMP, -I forces ICMP, -T -p 443 uses TCP SYNs so the probes follow whatever firewalls allow for the real traffic.
$ traceroute -n 203.0.113.5
traceroute to 203.0.113.5 (203.0.113.5), 30 hops max, 60 byte packets
1 10.1.1.1 0.412 ms 0.388 ms 0.371 ms
2 10.0.0.9 1.102 ms 1.080 ms 1.055 ms
3 * * *
4 198.51.100.14 8.913 ms 8.870 ms 8.851 ms
5 203.0.113.5 9.204 ms 9.180 ms 9.150 ms
$ mtr -rwzc 100 -n -T -P 443 203.0.113.5
HOST: client01 Loss% Snt Last Avg Best Wrst StDev
1. AS??? 10.1.1.1 0.0% 100 0.4 0.4 0.3 0.9 0.1
2. AS64500 10.0.0.9 0.0% 100 1.1 1.1 1.0 2.3 0.2
3. AS??? ??? 100.0% 100 0.0 0.0 0.0 0.0 0.0
4. AS64501 198.51.100.14 40.0% 100 8.9 9.0 8.8 11.2 0.4
5. AS64501 198.51.100.22 3.0% 100 9.1 14.8 9.0 88.0 15.1
6. AS64502 203.0.113.5 3.0% 100 9.2 15.1 9.1 90.2 15.4
- Hop 3 stars: that router does not send Time Exceeded (rate limited, filtered, or MPLS hiding the hop). Hops 4 and 5 answer, so traffic passes through it fine. Stars are only a problem when every hop after them is also a star, and even then the destination may simply filter the probe type.
- Read mtr from the last hop backwards. The destination shows 3% loss and a Wrst of 90 ms. Hop 4 shows 40%: hops 5 and 6 do not, so hop 4 is rate-limiting its own ICMP and is not the problem. The first hop where the loss and jitter match the destination is hop 5. That is where to look (interface counters, utilisation).
-rreport mode,-wwide,-zAS numbers,-c 100probes,-nno DNS,-T -P 443TCP to the real port,-uUDP,-i 0.2interval.- Forward path only. The reply to each probe takes the return path, which you cannot see. Run mtr from the far end too.
- ECMP hashes on the 5-tuple. Classic traceroute changes the destination port every probe, so consecutive probes may cross different parallel links and the hop list can interleave two paths. Fix the ports (
mtrwith-Lfor the source port, ortraceroute --sport=) and sweep source ports to walk the members one by one. Used in topic 19 scenario 2.
Names: dig
$ dig www.example.com
;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 31337
;; flags: qr rd ra; QUERY: 1, ANSWER: 2, AUTHORITY: 0, ADDITIONAL: 1
;; QUESTION SECTION:
;www.example.com. IN A
;; ANSWER SECTION:
www.example.com. 300 IN CNAME edge.example.net.
edge.example.net. 30 IN A 203.0.113.5
;; Query time: 23 msec
;; SERVER: 10.1.1.53#53(10.1.1.53) (UDP)
$ dig +short www.example.com @8.8.8.8
edge.example.net.
203.0.113.5
$ dig +trace www.example.com | grep -E '^(www|example|com)'
| Field | Read it as |
|---|---|
status: NOERROR | The name exists. With ANSWER: 0 it exists but has no record of that type (asked AAAA, only A exists). |
status: NXDOMAIN | The authoritative server says the name does not exist. Typo, missing record, wrong search domain, or a broken zone. |
status: SERVFAIL | The resolver could not get an answer: authoritative servers unreachable, DNSSEC validation failure, lame delegation. The resolver is up; its upstream is the problem. |
status: REFUSED | Policy: that server will not answer you (wrong ACL, asking an authoritative server for recursion). |
;; connection timed out; no servers could be reached | You never reached the resolver. This is a network or resolver-down problem, not a DNS data problem. |
flags: aa | Authoritative answer. Only present when you ask the authoritative server directly. ra = recursion available, rd = recursion desired. |
| ANSWER vs AUTHORITY | ANSWER holds the records you asked for. AUTHORITY holds the NS records (or SOA on a negative answer) telling you who owns the zone. |
| TTL (300, 30) | Seconds the answer may be cached. A low TTL on the A record is a load balancer or CDN moving traffic. A stale cache during a change shows as "my dig and yours differ": compare with @server. |
+trace | Walks root → .com → example.com authoritative yourself, bypassing every cache. Use it when resolvers disagree. |
dig -x 203.0.113.5 does the PTR lookup. Remember that dig speaks only DNS: getent hosts www.example.com goes through nsswitch (/etc/hosts first, then DNS) the way an application does. Resolution chain detail is in topic 6.
HTTP(S): curl -v and -w
$ curl -sv -o /dev/null https://www.example.com/ 2>&1 | grep -E '^(\*|>|<) ' | head -20
* Host www.example.com:443 was resolved.
* IPv4: 203.0.113.5
* Trying 203.0.113.5:443...
* Connected to www.example.com (203.0.113.5) port 443
* ALPN: curl offers h2,http/1.1
* TLSv1.3 (OUT), TLS handshake, Client hello (1):
* TLSv1.3 (IN), TLS handshake, Server hello (2):
* TLSv1.3 (IN), TLS handshake, Certificate (11):
* SSL connection using TLSv1.3 / TLS_AES_256_GCM_SHA384
* Server certificate:
* subject: CN=www.example.com
* SSL certificate verify ok.
* using HTTP/2
> GET / HTTP/2
> Host: www.example.com
< HTTP/2 200
< content-type: text/html
$ curl -s -o /dev/null -w 'dns %{time_namelookup} tcp %{time_connect} tls %{time_appconnect} ttfb %{time_starttransfer} total %{time_total} code %{http_code}\n' https://www.example.com/
dns 0.021 tcp 0.024 tls 0.061 ttfb 0.412 total 0.418 code 200
- The
-wtimes are cumulative from the start. Subtract: DNS = 21 ms; TCP handshake = 24 − 21 = 3 ms (one RTT); TLS = 61 − 24 = 37 ms (one RTT in TLS 1.3, two in 1.2); server think time = 412 − 61 = 351 ms; body transfer = 6 ms. This request is slow because of the server, not the network. - Each verbose stage that stops is a layer: no "resolved" = DNS; stuck at "Trying" = TCP (SYN dropped); stuck after "Client hello" = TLS blocked or a PMTUD black hole eating the large Certificate message; certificate verify failed = wrong name, expired, untrusted CA (corporate proxy), or a wrong client clock;
< HTTP/2 502= the edge reached, the backend did not. -kskips certificate verification: a diagnostic to separate "TLS is broken" from "the cert is wrong", never a fix.-Isends HEAD.-4/-6force the address family (a broken AAAA record hides behind happy eyeballs).--resolve www.example.com:443:203.0.113.9sends the request to one specific backend with the correct Host header and SNI, bypassing DNS and the load balancer. That is how you test "is it one backend or all of them".-H 'Host: x'alone does not set SNI.--connect-timeout 3 -m 10keeps a hung test from hanging you.-Lfollows redirects; without it you see the 301 itself.
The wire: tcpdump
$ sudo tcpdump -ni eth0 -c 5 'host 10.9.5.5 and tcp port 5432'
listening on eth0, link-type EN10MB (Ethernet), snapshot length 262144 bytes
12:00:01.000112 IP 10.1.1.10.40122 > 10.9.5.5.5432: Flags [S], seq 1998456112, win 64240, options [mss 1460,sackOK,TS val 1 ecr 0,nop,wscale 7], length 0
12:00:02.021330 IP 10.1.1.10.40122 > 10.9.5.5.5432: Flags [S], seq 1998456112, win 64240, options [mss 1460,sackOK,TS val 1021 ecr 0,nop,wscale 7], length 0
12:00:04.069420 IP 10.1.1.10.40122 > 10.9.5.5.5432: Flags [S], seq 1998456112, win 64240, options [mss 1460,sackOK,TS val 3069 ecr 0,nop,wscale 7], length 0
$ sudo tcpdump -ni any -s0 -w /tmp/case.pcap 'host 10.9.5.5 and not port 22'
$ sudo tcpdump -nr /tmp/case.pcap 'tcp[tcpflags] & (tcp-syn|tcp-rst) != 0'
| Flag | Meaning | Flag | Meaning |
|---|---|---|---|
[S] | SYN | [S.] | SYN-ACK (the dot is ACK) |
[.] | ACK only | [P.] | PSH-ACK, data being delivered |
[F.] | FIN-ACK, polite close | [R] / [R.] | RST: closed port, firewall reject, or abort |
- The capture above is the SYN-SENT picture: the same SYN three times at 1 s, 2 s, 4 s (exponential backoff), no
[S.]and no[R]. The packet left this host. Where it died needs a capture at the other end. - Options:
-nno name resolution (no stalls, no noise),-i eth0or-i any,-c Nstop after N,-s0whole packets (modern default is 262144 anyway),-w filewrite a pcap for Wireshark and-r fileread it back,-eshow MACs (which gateway did we actually send to),-v/-vvfor TTL, id and checksums,-A/-Xpayload. - Filters are BPF:
host,src host,dst port,net 10.9.0.0/16,icmp,arp,vlan,not port 22(always, when you are on SSH), and the flag test shown above for SYN/RST only. Quote the whole filter. - Where tcpdump sits: inbound packets are captured before netfilter, outbound after. So a SYN that shows up in tcpdump on the server but gets no SYN-ACK and never reaches the listener was dropped by the server's own firewall. And an outbound packet dropped by a local OUTPUT rule never appears in tcpdump at all.
- Offloads: with TSO/GRO you may see "segments" bigger than the MSS and wrong checksums on outbound packets. That is the NIC doing the work later, not corruption.
- Both ends: capture on the client and the server at the same time, filtering on the 4-tuple. Left A and arrived at B: the network is fine, look at B. Left A, never arrived at B: the drop is between them, split again at a midpoint. Arrived at B, B replied, A never saw the reply: return path. Arrived at B, no reply: B's firewall, no listener, or rp_filter.
Counters: ip -s link, ethtool -S
$ ip -s link show eth0
2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc mq state UP mode DEFAULT group default qlen 1000
RX: bytes packets errors dropped missed mcast
9812345678 8123456 1532 12044 0 4410
TX: bytes packets errors dropped carrier collsns
4123456789 5123456 0 0 0 0
$ ethtool -S eth0 | grep -Ei 'err|drop|miss|crc|over|fifo' | grep -v ': 0$'
rx_crc_errors: 1532
rx_missed_errors: 12044
rx_fifo_errors: 12044
$ ethtool eth0 | grep -E 'Speed|Duplex|Link detected'
Speed: 10000Mb/s
Duplex: Full
Link detected: yes
- errors / CRC / frame: bad bits arrived. Physical layer: cable, dirty or failing optic, bad connector, duplex mismatch. The link is up and lying.
- dropped / missed / fifo / overruns: good frames arrived and the host could not take them off the NIC ring in time. CPU, interrupt affinity, ring size (
ethtool -g), a softirq storm. Not the cable. - carrier on TX: the link flapped.
dmesg -T | grep eth0shows "Link is Down / Up" timestamps. - Counters are since boot. Take two readings a minute apart (
watch -d -n1 'ip -s link show eth0'). A million errors from last year is history; ten a second is the problem.
One-liners you should be able to say
| Tool | Line | Reading |
|---|---|---|
nc | nc -zv 10.9.5.5 5432 | "succeeded" = handshake done. "Connection refused" = RST came back: reachable, nothing listening. Hang then timeout = SYN dropped or no route. UDP (-u) can only prove "closed" when an ICMP port unreachable comes back. |
arping | arping -I eth0 -c 3 10.1.1.20, arping -D -I eth0 10.1.1.10 | L2 reachability even when ICMP is filtered. -D is duplicate address detection: any reply means someone else has your IP. Two MACs replying to the first form means a duplicate on the segment. |
conntrack | conntrack -L | grep 10.9.5.5, conntrack -S | The kernel's flow table behind NAT and stateful rules: state per flow, and insert_failed/drop counters. "nf_conntrack: table full, dropping packet" in dmesg explains new connections failing while old ones work. |
nft / iptables | nft list ruleset, iptables -S, iptables -L -n -v | Rules in order and the chain policy (ACCEPT or DROP). -v shows packet counters per rule: retry the failing connection and watch which counter moves. Also check sysctl net.ipv4.conf.all.rp_filter if packets arrive with "wrong" source routes. |
/etc/resolv.conf | cat /etc/resolv.conf | nameserver 127.0.0.53 is the systemd-resolved stub; the real upstream is in resolvectl status. search adds suffixes to short names. options timeout:2 attempts:2 bounds how long a dead resolver can hurt you. |
/etc/hosts | grep name /etc/hosts, getent hosts name | Wins over DNS when /etc/nsswitch.conf says hosts: files dns. A stale entry here makes dig and the application disagree. |
systemd-resolve | systemd-resolve --status (older) or resolvectl status, resolvectl query name, resolvectl flush-caches | Per-link DNS servers and search domains, which link a query will use, and a cache you can flush to rule out stale answers. |
Linux host health in one pass
NPE interviews are not the PE Linux round, but "what causes high CPU" and "how do you find a Python memory leak" have been asked. Have one command and one sentence per resource.
CPU
$ uptime
14:02:11 up 41 days, 3:12, 2 users, load average: 9.84, 6.12, 2.40
$ nproc
8
$ ps -eo pcpu,pmem,pid,etime,cmd --sort=-pcpu | head -4
%CPU %MEM PID ELAPSED CMD
98.2 1.1 2210 3-04:11:02 /usr/bin/python3 collector.py
3.0 0.4 1407 41-03:12 nginx: worker process
0.4 0.1 812 41-03:12 sshd: /usr/sbin/sshd
- Load average is the average number of runnable plus uninterruptible (disk-wait) tasks over 1, 5 and 15 minutes. Compare it to core count: 9.84 on 8 cores and rising from 2.40 means the box is saturating now, not historically.
- In
top, the%Cpu(s)line splits it: us user code (an application is busy), sy kernel (syscalls, network stack), wa waiting on disk, si softirq (packet processing pinned to one core on a busy NIC), st steal (a noisy neighbour VM). Press1to see per-core: one core at 100% si is a NIC queue problem, not an app problem. ps%CPU is the lifetime average of the process, so a long-lived process can hide a burst. Usetop -p PIDorpidstat 1for now. Thenstrace -p PID -corpy-spy dump --pid PIDto see what the code is doing.- Usual causes: a busy loop or retry storm, GC or a regex on a huge string, swap thrashing (looks like CPU but is
wa), interrupt storms from a NIC, a cron job overlapping itself.
Memory
$ free -m
total used free shared buff/cache available
Mem: 32000 21500 900 120 9600 10100
Swap: 2047 800 1247
$ dmesg -T | grep -i 'killed process'
[Tue ...] Out of memory: Killed process 2210 (python3) total-vm:9123456kB, anon-rss:8234560kB
- Read available, not free. Linux fills spare RAM with page cache (
buff/cache) on purpose and gives it back on demand. 900 MB free with 10 GB available is a healthy box. - Swap in use is not a crisis; swap activity is:
vmstat 1columnssi/sonon-zero means pages moving right now. - The OOM killer picks the process with the highest oom_score (mostly RSS) and logs it in the kernel ring buffer.
dmesg -Torjournalctl -kis where "the service just vanished" is answered.
Finding a Python memory leak
$ while true; do ps -o rss= -p 2210; sleep 60; done # RSS in kB, does it plateau or climb forever?
import tracemalloc
tracemalloc.start(25)
s1 = tracemalloc.take_snapshot()
# ... run the workload for a while ...
s2 = tracemalloc.take_snapshot()
for stat in s2.compare_to(s1, 'lineno')[:5]:
print(stat)
# collector.py:88: size=412 MiB (+398 MiB), count=3100000 (+3000000), average=139 B
- First prove it is a leak: sample RSS over time. A cache that fills and plateaus is not a leak. Monotonic growth across hours is.
tracemalloccompares two snapshots and names the line allocating the growth.objgraph.show_growth()names the object types growing between calls (3 million dicts appearing says "something is appending to a list of dicts").- Usual culprits: a module-level list or dict that only grows,
functools.lru_cachewithoutmaxsize, references held by closures, logging handlers or exception tracebacks, a C extension leaking (tracemalloc cannot see it), or allocator fragmentation where freed memory is never returned to the OS (RSS stays high, tracemalloc shows nothing). - Mitigation while you fix it: bound the cache, restart on an RSS threshold, or run under a supervisor with a memory limit.
Disk, logs, file descriptors
$ df -h / # Use% 100%: "No space left on device"
$ df -i / # inodes can fill while bytes are free
$ du -xsh /var/log/* 2>/dev/null | sort -rh | head
$ lsof +L1 # deleted files still held open: space not freed until the process closes them
$ journalctl -u nginx -S -1h -p err --no-pager
$ journalctl -k -S -30min # kernel: link up/down, OOM, conntrack full
$ dmesg -T | tail -50
$ ls /proc/2210/fd | wc -l ; grep 'open files' /proc/2210/limits
$ lsof -p 2210 | awk '{print $5}' | sort | uniq -c | sort -rn | head
$ ulimit -n # this shell only; services get LimitNOFILE= from systemd
- A full disk breaks logging, DNS caches, TLS session files and databases in confusing ways. Truncate the big log (
: > file) rather thanrmit, because a deleted file that is still open keeps its blocks. journalctl -u unitis the service's view;-k/dmesg -Tis the kernel's view. "NETDEV WATCHDOG", "Link is Down", "nf_conntrack: table full" and "Out of memory" all live in the kernel log.- "Too many open files" (EMFILE) is a per-process limit, 1024 by default in many shells. Count the fds in
/proc/PID/fd, compare with/proc/PID/limits. Thousands of sockets in CLOSE-WAIT inssplus this error means the application never callsclose().
Interview questions
1. What does a star in traceroute mean?
No reply came back for that probe within the timeout. The usual reason is that the router at that hop does not generate ICMP Time Exceeded, rate-limits it, or a filter drops it; MPLS cores often hide hops this way. If later hops answer, traffic is flowing through the starred hop and it is not a problem. Only stars that continue to the end suggest a real break, and even then the destination may just be filtering the probe type, so retry with-I or -T -p 443.2. ss shows a socket in SYN-SENT. What does that tell you?
The host sent a SYN and received neither a SYN-ACK nor a RST. So the far end is silent to us: no route, a silent firewall drop in either direction, the server down, or its reply lost on the return path. It rules out "port closed", because a closed port would answer RST and the socket would be gone. Next step: tcpdump on the client to see the SYN retransmits leave, then a capture on the server to see whether they arrive.3. How would you confirm a packet actually left the host?
tcpdump -ni eth0 filtered on the destination, while generating the traffic. If the packet appears on the egress interface, it left. Two caveats: tcpdump sees outbound packets after netfilter, so a local OUTPUT drop means it never appears, and seeing it leave proves nothing about arrival. Pair it with ip route get to confirm it went out the interface you expected, and -e to check the destination MAC is the right gateway.4. What is the difference between UP and LOWER_UP on an interface?
UP is the administrative state: the interface is enabled in software. LOWER_UP is the carrier: the NIC detects a live link on the cable. UP without LOWER_UP (shown as NO-CARRIER) is a physical problem or a down switch port; LOWER_UP without an address or route is a configuration problem.5. What do Recv-Q and Send-Q mean?
On a LISTEN socket, Recv-Q is the number of completed connections waiting for accept() and Send-Q is the backlog limit; Recv-Q near Send-Q means the app is not accepting and SYNs are being dropped. On an established socket, Recv-Q is bytes received that the application has not read (slow app) and Send-Q is bytes sent but not yet acknowledged (slow network or peer).6. dig returns NXDOMAIN, SERVFAIL, or times out. What does each mean?
NXDOMAIN: the authoritative server says the name does not exist, so look at the zone or the spelling. SERVFAIL: the resolver tried and failed, usually because the authoritative servers are unreachable or DNSSEC validation failed. Timeout: you never reached the resolver at all, which is a network or resolver-down problem. Follow-up: ask the same question of a different resolver with@8.8.8.8 and of the authoritative server with +trace to see who disagrees.7. How do you find the path MTU with ping?
Send pings with the DF bit set (-M do) and a payload size, starting at 1472 (1500 minus 28 bytes of headers). If a link on the path is smaller, the router returns ICMP Fragmentation Needed with its MTU, and you step down. If large probes silently disappear while small ones work, something is dropping that ICMP, which is the PMTUD black hole from topic 8.8. mtr shows 40% loss at hop 4 and 0% at the destination. Is there a problem?
No. Loss at an intermediate hop that does not continue to later hops is that router rate-limiting its own ICMP replies; transit traffic through it is fine. Real loss starts at some hop and persists to the final hop, because every later probe has to cross the bad link. Read the table from the bottom up.9. A page takes two seconds. How do you tell DNS from network from server?
curl -w with time_namelookup, time_connect, time_appconnect, time_starttransfer and time_total, then subtract adjacent values. DNS is the first number. Connect minus namelookup is one RTT, so a value of one or three seconds is a lost SYN. TLS is appconnect minus connect, one RTT in TLS 1.3. Starttransfer minus appconnect is the server's think time. The remainder is transfer, where loss and window size show up. Each bucket points to a different team.10. Load average is 12 on an 8-core server. What does that mean, and what do you type next?
On average 12 tasks were runnable or stuck in uninterruptible I/O over the last minute, so the box is oversubscribed. Compare the 1, 5 and 15 minute numbers to see if it is rising. Thentop for the us/sy/wa/si split (user code vs kernel vs disk vs softirq) and ps --sort=-pcpu or pidstat to name the process. High wa with low us means it is a disk problem wearing a CPU costume.Traps
- Treating ping as proof either way. A reply proves ICMP works right now; silence proves only that ICMP did not come back. Test the real port with
ncorcurl. - Reading per-hop traceroute or mtr loss as real loss. Only loss that persists to the destination counts.
- Forgetting
-n. Reverse DNS lookups in traceroute and tcpdump add seconds that look like latency and stall the capture. - Saying "tcpdump shows nothing, so the packet never arrived" on the sender. Outbound drops by the local firewall are invisible to tcpdump; use
iptables -L -n -vcounters. - Reading the
freecolumn infree -mand declaring memory exhausted. Readavailable. - Quoting
ps%CPU as current usage. It is a lifetime average. - Trusting dig to describe what the application sees. The application goes through
/etc/hostsand nsswitch; dig does not. - Reading a counter once. Errors and drops are cumulative since boot; the rate is the finding.
Scenario
A health check against a web server started failing ten minutes ago. You have a shell on the server and nothing else. Which commands, in which order, and what does each one rule out?
Expected reasoning
ss -tlnp | grep ':443': is anything listening, and on which address? Nothing listening means the service died:systemctl status,journalctl -u nginx -S -15min,dmesg -T | grep -i killedfor the OOM killer. Listening on 127.0.0.1 only means a config change broke the bind address.curl -sv -o /dev/null https://localhost/ --resolve www.example.com:443:127.0.0.1: does the service answer locally? A 200 here moves the fault off the application and onto the path in. A 502 or a hang keeps it on the box (backend, disk full, CPU).ip -br link,ip addr,ip route get <health-checker IP>: link up with carrier, the right address and mask, a route back to the checker. No route back is a complete explanation on its own.tcpdump -ni eth0 'host <checker> and port 443'while the check runs. SYNs arrive and SYN-ACKs leave: the server is fine, the return path or the checker is the problem. SYNs arrive, no SYN-ACK: local firewall (nft list ruleset, look for a new DROP, check counters) or the listener is gone. Nothing arrives: upstream, so escalate with "the box is healthy and sees no traffic".uptime,free -m,df -htake ten seconds and catch the boring causes: a disk that filled ten minutes ago, a load spike, an OOM.- Report in one line: what is confirmed, what is ruled out, what you are testing next. The structure is the answer; the commands are the evidence.
← 17 · Classics: binary search, recursion, trees and BST · all topics · 19 · Troubleshooting walkthroughs →