← all topics

12 · Routing fundamentals, OSPF (vs BGP) and IS-IS Network

RIB vs FIB · longest prefix, admin distance, metric · link-state: hellos, LSDB, SPF · areas and scale limits · why OSPF can't run the whole Internet · IS-IS in a paragraph

Why it matters for NPE. 8 of 22 reports asked OSPF, almost always as 'why BGP instead of OSPF' or 'explain an IGP in depth'. You need one IGP cold and the comparison fluent. IS-IS is a preferred qualification in the posting, so know what it is.

Primer: a router is a lookup table with opinions about how to fill it

Every router does two separate jobs. The control plane learns routes (static config, OSPF, IS-IS, BGP) and argues about which one is best. The data plane looks up each packet's destination, rewrites the layer-2 header and sends it out an interface, millions of times a second, without thinking. An IGP (interior gateway protocol) is the control-plane protocol that one organisation runs inside its own network, where every router is trusted and the goal is simply the shortest path. OSPF and IS-IS are the two link-state IGPs; both give every router a complete map and let it compute its own shortest-path tree. BGP, covered in topic 14, is the protocol between organisations where trust and policy replace "shortest".

The sentence to open with: "OSPF is a link-state IGP. Routers discover neighbours with hellos, flood link-state advertisements so every router in the area holds the same database, and each one runs Dijkstra over that database to build its own shortest-path tree. It converges in seconds, trusts every router, and scales by splitting the network into areas."

Watch

Free CCNA | OSPF Part 1 | Day 26 | CCNA 200-301 Complete CourseJeremy's IT Lab · 39:40

Watch 2:03-18:44: link-state in plain terms (2:43), LSA flooding (5:41), the three-step process (8:25), areas (9:08) and area rules (15:52). Stop at 18:44; the rest is Cisco configuration and quizzes.

Free CCNA | OSPF Part 2 | Day 27 | CCNA 200-301 Complete CourseJeremy's IT Lab · 36:55

Watch 1:39-22:11: cost and reference bandwidth (1:39, 4:01), then every neighbour state with a packet diagram (11:06 Down through 18:53 Full) and the message-type chart at 22:11. This is the part interviewers probe.

Route Precedence -- How does a Router choose a path when multiple paths exist?Practical Networking · 15:00

The tie-break order in a lab: ECMP (1:56), metric (4:35), administrative distance (8:10), longest prefix (11:54). Fifteen minutes that fix the order in your head permanently.

Free CCNA | Routing Fundamentals | Day 11 (part 1) | CCNA 200-301 Complete CourseJeremy's IT Lab · 31:00

Only if the routing table itself is new to you: 7:54-21:01 covers connected and local routes and longest-prefix selection.

Intermediate System to Intermediate System (IS-IS) Routing Protocol FundamentalsKevin Wallace Training, LLC · 14:58

Levels 1 and 2, NET addresses, TLVs, why operators like it. Watch whole; it covers everything the IS-IS section below says, with pictures.

OSPF and IS-IS: A Comparative AnatomyNANOG · 50:48

Optional depth from Dave Katz, co-author of both protocols. The first 25 minutes on protocol structure are the useful part.

Comparing IS-IS and OSPFNSRC Network Startup Resource Center · 4:05

Four-minute summary if you only need the comparison table spoken aloud.

Reading: RFC 7938 §7 gives the convergence vocabulary (detection, propagation, SPF, FIB update) you want in the "how fast does it converge" answer.

RIB, FIB and the two planes

RIB (routing table)FIB (forwarding table)
Lives incontrol plane, CPU memory, one per routing process plus the merged tabledata plane, the ASIC or kernel forwarding path
Containsevery candidate route from every source with its metric and sourceonly the winners: prefix, next hop, egress interface, rewritten MAC (the adjacency)
Built byprotocols and static configpushed from the RIB after best-route selection; the next hop is resolved to a layer-2 adjacency (ARP/ND)
Linuxip route show table all, FRR show ip routethe kernel's route cache / fib trie; ip route get 10.2.2.5 shows the lookup result

Why this matters: a route can be in the RIB and still not forward (next hop unresolvable, interface down, FIB full). Hardware FIBs have a fixed size; that is why the Internet's million prefixes do not fit on a cheap switch and why Meta summarises aggressively inside its fabrics.

How one route is chosen

  1. Longest prefix match. For a given packet the most specific route wins. A /32 beats a /24 beats 0.0.0.0/0 whatever their source. This is a per-packet lookup rule, not a RIB rule.
  2. Administrative distance (Cisco) or route preference (Juniper). When two sources offer the same prefix, the more trusted source is installed in the RIB.
  3. Metric. Within one protocol, lowest metric wins (OSPF cost, IS-IS metric, RIP hops, BGP's attribute chain).
  4. ECMP. Equal-metric routes from the same protocol are all installed and traffic is hashed across them (OSPF: 4 by default on Cisco, up to 64 or 128 on modern platforms).
SourceCisco ADJuniper preferenceFRR distance
Connected000
Static151
eBGP20170 (all BGP)20
OSPF (internal)11010 internal, 150 external110
IS-IS11515 L1, 18 L2115
RIP120100120
iBGP200170200

Two things to say out loud: the numbers are vendor conventions, and Juniper prefers OSPF over eBGP while Cisco does the opposite. The plain Linux kernel has no administrative distance at all: it keeps one route per prefix per metric and a protocol tag (proto bgp, proto ospf), so the routing daemon (FRR, BIRD, Meta's own agent) must do the comparison before it installs. Multiple routing tables plus policy rules (ip rule) give Linux policy routing that Cisco does with route-maps.

Static routes, default, floating, recursive

ip route 10.9.0.0/16 via 192.0.2.1            # static: next hop on a connected subnet
ip route 0.0.0.0/0  via 192.0.2.1             # default route: the /0 that matches when nothing else does
ip route 0.0.0.0/0  via 198.51.100.1 metric 250   # floating static: higher cost, installed only when the primary disappears
ip route 172.16.0.0/12 via 10.9.1.1           # recursive: 10.9.1.1 is not connected, so the router
                                              # looks up 10.9.1.1 (hits 10.9.0.0/16) to find the real egress

A floating static on Cisco is a static with a higher AD than the dynamic route (ip route 0.0.0.0 0.0.0.0 198.51.100.1 250); it floats in only if OSPF or BGP stops offering that prefix. A recursive next hop that points at a prefix learned by BGP gives you the classic "route is there but traffic blackholes" when the inner route flaps. Static routes never go away on their own unless the interface drops, which is why operators pair them with BFD or IP SLA tracking.

Distance-vector vs link-state

Distance-vector (RIP, EIGRP's hybrid)Link-state (OSPF, IS-IS)
What a router knowsits neighbours' distances to each prefix ("routing by rumour")the whole topology of its area: every router, every link, every cost
What it sendsits routing table to neighbours, periodicallyits own links (LSAs), flooded once to everybody, refreshed every 30 min
AlgorithmBellman-FordDijkstra SPF, each router computes independently
Loop preventionsplit horizon, poison reverse, hold-down, max hop counta consistent database gives a loop-free tree; only transient micro-loops during convergence
Convergenceslow (RIP: 30 s updates, 180 s timeout)fast (sub-second detection with BFD, SPF in ms)
Costcheap CPU and memorymemory for the LSDB, CPU for SPF, flooding traffic on change
Path-vector (BGP)distance-vector that carries the full AS path instead of a number: loop-free by path inspection, policy on every hop, no topology view

OSPF in depth

Packets and addresses

OSPF rides directly in IP as protocol 89, no TCP or UDP. On broadcast segments it multicasts to 224.0.0.5 (AllSPFRouters) and 224.0.0.6 (AllDRouters, listened to by the DR and BDR). OSPFv3 for IPv6 uses ff02::5 and ff02::6. Five packet types:

TypeNameJob
1HelloDiscover and keep neighbours. Carries router ID, area ID, netmask, hello and dead intervals, authentication, options (E-bit for externals), priority, current DR/BDR, and the list of neighbours already seen.
2DBD (database description)Summary of my LSDB: LSA headers only. Also negotiates master/slave and carries the interface MTU.
3LSR (link-state request)"Send me the full copy of these LSAs I am missing or that are newer than mine."
4LSU (link-state update)The LSAs themselves. Also used for flooding on change.
5LSAckReliable flooding: every LSU is acknowledged or retransmitted.

Neighbour states

Down ──hello rcvd──▶ Init ──my RID in its hello──▶ 2-Way ──(DR/BDR or p2p)──▶ ExStart ──master/slave agreed──▶ Exchange ──DBDs done──▶ Loading ──LSR/LSU──▶ Full │ └── DROther to DROther stays here forever (correct, not a fault)

What must match for an adjacency

Must matchSymptom when it does not
Area ID on the shared linkhello discarded; neighbour never leaves Down (log: "mismatched area")
Hello and dead timers (10/40 s on broadcast and point-to-point, 30/120 s on NBMA)hello discarded; stuck Down
Subnet and mask (except on point-to-point)hello discarded
Authentication type and key (none, plain, MD5, or SHA via RFC 5709)hello discarded
Stub / area-type flags (the E-bit and N-bit in options)hello discarded
MTUreaches ExStart, loops there
Network type (broadcast vs point-to-point)adjacency may reach Full, but the LSAs describe the link differently and routes across it are missing. The sneakiest one.
Unique router IDsduplicate RIDs make routers fight over the same LSAs; the LSDB churns and routes flap

Hellos do not need matching priority, cost or router ID. Cost mismatches only make paths asymmetric.

DR and BDR on broadcast segments

On a multi-access segment with n routers, full adjacencies between all pairs would mean n(n-1)/2 flooding relationships and n Network LSAs' worth of mess. OSPF elects a Designated Router and a Backup DR; every other router (a DROther) forms a Full adjacency only with the DR and BDR and stays in 2-Way with the rest. DROthers send updates to 224.0.0.6, the DR re-floods to 224.0.0.5. Election: highest priority (default 1, 0 means never), tie broken by highest router ID. The election is not preemptive: a new router with a higher priority waits until the current DR goes away. The DR generates the type 2 Network LSA for the segment. Point-to-point links skip all of this, which is why fabric links are configured ip ospf network point-to-point.

LSA types

TypeNameOriginated byScopeDescribesShows as
1Routerevery routerits areamy router ID, my links, their costs and statesO
2Networkthe DRits areathe routers attached to a broadcast segment and its maskO
3SummaryABRother areasa prefix from another area with its cost (optionally summarised)O IA
4ASBR summaryABRother areashow to reach an ASBR that sits in a different area(enables E routes)
5AS externalASBRwhole domain except stub areasa prefix redistributed from outside OSPF (BGP, static). E1 adds internal cost, E2 (default) keeps only the external metricO E1 / O E2
7NSSA externalASBR inside an NSSAthat NSSA; the ABR translates to type 5externals in a not-so-stubby areaO N1 / O N2

Types 1 and 2 are the topology. Types 3, 4, 5 and 7 are reachability only, which is the point: inter-area and external routes are carried distance-vector style so the SPF tree is computed per area.

LSDB, SPF and cost

Every router in an area holds the same link-state database (LSDB): the set of all type 1 and 2 LSAs for that area plus the 3/4/5s it has received. On any change it runs Dijkstra with itself as the root to build a shortest-path tree, then installs the leaves into the RIB. SPF cost is O(E log V), so a few thousand nodes is fine and a hundred thousand is not. Each interface has a cost = reference bandwidth / interface bandwidth, with reference bandwidth 100 Mbit/s by default: 10 Mbit = 10, 100 Mbit = 1, and everything from 1 Gbit upward also rounds to 1, so a 1G link and a 100G link look identical until you raise the reference (auto-cost reference-bandwidth 100000, identically on every router). Cost is cumulative along the path and asymmetric costs are allowed. IS-IS does not care about bandwidth at all: every interface is 10 unless you set it.

Areas, ABR and ASBR

area 1 area 0 (backbone) area 2 R1 ── R2 ── [ABR-A] ───────── R5 ── R6 ── [ABR-B] ───── R8 ── R9 │ [ASBR] ── eBGP to the ISP (type 5 LSAs) Rules: every area touches area 0; inter-area traffic crosses area 0; summarise at the ABR.

Areas exist because flooding and SPF do not scale: a flapping link anywhere in an area makes every router in that area re-flood and recompute. An ABR (area border router) sits in two or more areas, holds an LSDB per area, and injects type 3 summaries in each direction; it is the only place you can summarise intra-OSPF prefixes. An ASBR redistributes external routes in as type 5. The backbone, area 0, must be contiguous and every other area must attach to it (or be stitched with a virtual link, which is a hack to admit). Area types in one line each: a stub area blocks type 5 and gets a default from the ABR; a totally stubby area also blocks type 3; an NSSA is a stub that is still allowed its own ASBR via type 7. The practical limit is a few hundred routers per area and a few hundred areas; beyond that is where BGP starts.

Convergence, step by step

  1. Detect: carrier loss (instant), BFD (tens of ms) or the dead timer (40 s by default; that is the number to quote and then say "which is why you tune it or run BFD").
  2. Originate: the routers on each side issue a new type 1 (and the DR a new type 2) with a higher sequence number.
  3. Flood: LSUs to every neighbour, acknowledged, hop by hop, through the area. Milliseconds across a campus.
  4. SPF: each router waits its SPF throttle (initial 50 ms on modern defaults, backing off under churn) and recomputes. Partial or incremental SPF for type 3/5 changes.
  5. RIB then FIB: new next hops installed, ASIC programmed. Total: well under a second with BFD, 40 seconds or more without it.

During steps 3-5 different routers hold different views for a few ms, so micro-loops are possible; that is the honest answer to "is link-state always loop-free".

What you'd type and see

! Cisco IOS-style
router ospf 1
 router-id 10.0.0.1
 auto-cost reference-bandwidth 100000     ! 100 Gbit/s = 1, must match everywhere
 network 10.0.0.0 0.0.255.255 area 0
 passive-interface default
 no passive-interface GigabitEthernet0/1
interface GigabitEthernet0/1
 ip ospf network point-to-point
 ip ospf hello-interval 1
 ip ospf dead-interval 3

R1# show ip ospf neighbor
Neighbor ID     Pri   State           Dead Time   Address         Interface
10.0.0.2          1   FULL/  -        00:00:02    10.0.12.2       GigabitEthernet0/1   ! "-" = p2p, no DR
10.0.0.3          1   FULL/DR         00:00:35    10.0.13.3       GigabitEthernet0/2
10.0.0.4          1   2WAY/DROTHER    00:00:38    10.0.13.4       GigabitEthernet0/2   ! normal on a LAN

R1# show ip route ospf
O     10.0.23.0/30 [110/2] via 10.0.12.2, 00:04:11, GigabitEthernet0/1      ! [AD/cost]
O IA  10.1.0.0/16  [110/3] via 10.0.12.2, 00:04:11, GigabitEthernet0/1
O E2  0.0.0.0/0    [110/1] via 10.0.13.3, 00:02:50, GigabitEthernet0/2

# FRR on Linux
vtysh -c 'show ip ospf neighbor'
vtysh -c 'show ip ospf database'
tcpdump -ni eth0 'ip proto 89'          # hellos every 10 s to 224.0.0.5

Why OSPF cannot run the whole Internet (and why not everywhere)

This is the question behind half the OSPF questions in the reports. Give four reasons, then the comparison.

Inside a large data center the same argument applies at smaller scale (see topic 16 and RFC 7938): thousands of switches make a single LSDB fat and flooding noisy, and operators want per-link policy and summarisation that OSPF's area model gives awkwardly. eBGP between every tier, with a private ASN per tier, is simpler to reason about and to write software for.

OSPFBGP
Kindlink-state IGPpath-vector EGP
TransportIP protocol 89, multicast hellos, discovers neighboursTCP 179, unicast, neighbours configured by hand
What it exchangestopology (LSAs); every router has the mapprefixes with attributes; nobody has the map
Best pathlowest cumulative cost, Dijkstraattribute comparison steered by policy
Loop preventionconsistent LSDB, SPF tree; area 0 hierarchy for inter-areaAS_PATH check (eBGP), split horizon and route reflector attributes (iBGP)
Convergencesub-second to secondsseconds to minutes (hold time 180 s default, MRAI)
Scalehundreds of routers per areathe Internet, a million prefixes
Trustevery routernobody; filter everything
Useinside an organisation, campus, small DC, reaching BGP next hopsbetween organisations, large DC fabrics, anywhere policy matters

IS-IS in a compact section

IS-IS (Intermediate System to Intermediate System, ISO 10589, RFC 1195 for IP) is the other link-state IGP and does the same thing as OSPF: hellos (called IIH), a flooded database of LSPs (link-state PDUs, one per router rather than many small LSAs), synchronised with CSNP/PSNP packets, and Dijkstra on every router. The differences are what interviewers want:

The one-line answer: "IS-IS is link-state like OSPF, but it runs on layer 2 instead of IP, uses levels instead of areas, and its TLV encoding made it easy to extend, which is why ISPs and backbones prefer it."

Interview questions

1. Choose an IGP and explain how it works in depth.OSPF. Routers send hellos on protocol 89 to 224.0.0.5 every 10 s; neighbours with matching area, timers, mask and authentication go through Init, 2-Way, ExStart, Exchange, Loading to Full, exchanging database descriptions and then the missing LSAs. Each router floods a type 1 LSA describing its links and costs; the DR on each broadcast segment adds a type 2. Every router in the area ends up with the identical LSDB and runs Dijkstra from itself, cost = reference bandwidth / link bandwidth summed along the path, and installs the tree into the RIB with AD 110. Areas bound flooding and SPF; ABRs carry type 3 summaries between areas through area 0; ASBRs inject externals as type 5. Follow-up you will get: "what happens when a link fails", so finish with the convergence steps.
2. Why can't we use OSPF to connect the whole Internet?Trust, policy, scale and boundaries. Link-state requires every router to trust and see every other; organisations will not share topology. OSPF has one metric and no per-neighbour policy, and the Internet is run on commercial policy. Flooding a million prefixes and every flap to every router with every router rerunning SPF does not scale, and areas only buy a factor of hundreds. BGP hides topology behind an AS number, applies policy per peer and is incremental over TCP.
3. Compare OSPF and BGP in one minute.OSPF is a link-state IGP over IP protocol 89 that floods LSAs so every router has the map and computes shortest path by cost; it converges in seconds, trusts everyone, and scales to hundreds of routers per area. BGP is a path-vector protocol over TCP 179 that exchanges prefixes with attributes and no topology; it picks paths by policy, converges slowly, trusts nobody and scales to the Internet. Use OSPF or IS-IS inside, BGP between organisations and in very large fabrics.
4. How do OSPF and BGP each avoid loops?OSPF: every router in an area computes a shortest-path tree from the same database, and a tree has no loops; between areas, summaries only cross area 0 so the inter-area graph is a star. The only loops are transient micro-loops while routers converge at slightly different times. BGP: eBGP rejects any route whose AS_PATH already contains the local AS; iBGP never re-advertises iBGP-learned routes, and route reflectors add ORIGINATOR_ID and CLUSTER_LIST to catch what that rule misses.
5. Two routers on the same link will not form an OSPF adjacency. Why?Ask what state they are stuck in. Never leaving Down: hellos are not being accepted, so check area ID, hello/dead timers, subnet mask, authentication, stub flags, and whether the interface is passive or an ACL blocks protocol 89. Stuck in Init: one side hears the other but not vice versa, usually a one-way filter or unicast/multicast problem. Stuck in ExStart or Exchange: MTU mismatch, or duplicate router IDs. Full but no routes: network-type mismatch. Confirm with show ip ospf neighbor, debug ip ospf adj or tcpdump ip proto 89.
6. What is the difference between the RIB and the FIB?The RIB is the control-plane table of every candidate route with its source and metric; best-route selection by prefix, administrative distance and metric picks winners. The FIB is the data-plane copy of only the winners, with the next hop resolved to an egress interface and MAC, programmed into the ASIC or kernel for per-packet lookups. A route in the RIB whose next hop will not resolve never reaches the FIB, which is the classic "route is there, traffic drops".
7. A router has an OSPF route and a static route to the same prefix. Which does it use, and what if the prefixes differ?Same prefix: administrative distance decides, so the static (AD 1) wins over OSPF (110) on Cisco; a floating static with AD 115 would lose. Different prefix lengths: longest prefix match decides per packet before AD is even considered, so an OSPF /24 beats a static /16 for destinations inside the /24.
8. What is the DR, why does it exist and how is it elected?On a multi-access segment OSPF elects a designated router and a backup so that every other router adjacency-pairs only with those two instead of with everyone, cutting flooding from n squared to 2n. The DR originates the type 2 network LSA for the segment. Election is highest priority then highest router ID, and it is not preemptive. Point-to-point links do not elect one.
9. Name the LSA types and who sends them.Type 1 router LSA from every router, type 2 network LSA from the DR, type 3 summary and type 4 ASBR-summary from ABRs, type 5 external from the ASBR, type 7 NSSA external from an ASBR inside an NSSA. Types 1 and 2 are topology within the area; the rest are reachability carried between areas.
10. How is OSPF cost calculated, and what goes wrong at 10 Gbit/s?Cost = reference bandwidth / interface bandwidth with a 100 Mbit/s reference by default, minimum 1. Anything 100 Mbit/s or faster is cost 1, so 1G, 10G and 100G links are indistinguishable and traffic may prefer a slow path. Fix by setting the same higher reference bandwidth on every router in the domain.
11. Why do areas exist and what are the rules?To limit the blast radius of flooding and SPF: a change inside one area only causes recomputation in that area, and ABRs can summarise. Rules: area 0 is the backbone, every other area must connect to it, inter-area traffic passes through it, and summaries (type 3) are not re-advertised from one non-backbone area into another, which keeps the inter-area topology loop-free.
12. What is IS-IS and why would a backbone prefer it to OSPF?A link-state IGP with the same hellos, flooding and Dijkstra as OSPF, but it runs directly on layer 2 instead of IP, addresses routers with a NET, uses Level 1 and Level 2 instead of areas with the boundary on a link, and encodes everything as TLVs. Backbones chose it for stability, large flat Level 2 domains, fewer moving parts and easy extension to IPv6, traffic engineering and segment routing.

Traps

Scenario

After a maintenance, a new switch was added to a broadcast segment in area 0. Its OSPF neighbours show EXSTART/DR toward the existing core router and never move. Users behind the new switch cannot reach anything beyond it. Walk through it.

Expected reasoning
  1. ExStart means hellos are fine (area, timers, mask, auth all match) and the problem is in the DBD exchange. The two classic causes are an MTU mismatch and duplicate router IDs.
  2. show ip ospf interface on both sides: compare MTU. A new switch with 9216 jumbo frames against a core at 1500 will loop here because the DBD advertises the MTU and the smaller side rejects it. Also compare router IDs; a cloned config often carries the same loopback.
  3. Fix: set matching MTU (or ip ospf mtu-ignore as a temporary measure and note it), or give the new switch a unique router ID and clear ip ospf process.
  4. Verify: state goes Exchange, Loading, Full; show ip ospf database shows the new type 1 LSA; show ip route ospf on the core shows the new prefixes with [110/cost]; ping end to end.
  5. Close the loop: why did it reach users? With no Full adjacency the new switch never got the LSDB, so it had no routes beyond its connected interfaces. Say that explicitly; the interviewer wants the symptom tied to the mechanism.

← 11 · Stacks, queues, intervals and 2D grids · all topics · 13 · Graphs: BFS and DFS on grids and adjacency lists →