← all topics

14 · BGP in depth Network

AS, eBGP vs iBGP, TCP 179, FSM, UPDATE attributes, best-path order, policy, route reflectors, loop prevention

Why it matters for NPE. BGP is the most-reported networking topic for Meta NPE. Pick it as your 'protocol I know best' and be ready for 15 minutes of follow-ups.

Primer: BGP is a policy protocol that happens to find paths

Inside one organisation an IGP (OSPF, IS-IS) finds the shortest path and trusts every router. Between organisations nobody trusts anybody and "shortest" is not the goal: money, contracts and traffic engineering are. BGP (Border Gateway Protocol, RFC 4271) is a path-vector protocol: each route carries the list of autonomous systems it has crossed, and every router applies policy to decide what to accept, what to prefer and what to tell its neighbours. It runs over TCP port 179, so it inherits reliability and needs no timers for retransmission of its own.

The sentence to open with: "BGP is a path-vector routing protocol between autonomous systems. It exchanges prefixes with attributes over a TCP session, prevents loops with the AS_PATH, and picks one best path per prefix with a strict attribute comparison that operators steer with policy."

Watch

BGP - Complete ENCOR (350-401) Exam CoverageKevin Wallace Training · 25:57

First pass: eBGP vs iBGP, how neighbours form (TCP 179, the states), the attribute list and best-path order, a basic eBGP config. Watch all of it, then read the tables below with the video paused.

BGP Deep DiveKevin Wallace Training · 2:10:27

Depth. Watch 19:46-36:16 (basics, confederations, route reflectors, neighbour formation) and 42:24-1:00:03 (path selection). Skip the synchronisation and lab demo sections.

Influencing BGP Path SelectionKevin Wallace Training · 16:47

Weight, local preference, AS-path prepending and MED with labs. Say "local preference", not "weight", when asked for a vendor-neutral knob.

NSDI '21 - Running BGP in Data Centers at ScaleUSENIX (Meta authors) · 11:23

Meta's own talk: why eBGP-only in the fabric, ASN per tier, summarisation, a tiny policy set, an in-house BGP agent tested like software. Watch whole; it is your "why Meta runs BGP in the DC" answer.

Tutorial: Introduction to BGPNANOG (Avi Freedman) · 1:35:30

Optional operator-level depth: policy with local-pref, MED and communities, route reflectors, when to use BGP instead of an IGP. The first hour is the core.

Reading that beats video here: RFC 4271 §9.1 (the decision process), RFC 7938 §5-6 (BGP in large data centers) and Meta's 2021 post on BGP in the data center.

Vocabulary

TermMeaning
AS (autonomous system)A network under one administration with one routing policy, identified by an ASN (16-bit, now 32-bit; 64512-65534 and 4200000000+ are private). Meta is AS32934.
Prefix / NLRIWhat BGP advertises: a network with its length (203.0.113.0/24) plus path attributes. NLRI = Network Layer Reachability Information.
eBGPA session between different ASes. Usually directly connected (TTL 1 by default). Prepends the local AS to AS_PATH and rewrites NEXT_HOP when advertising.
iBGPA session inside one AS, carrying external routes across it. Does not change AS_PATH or NEXT_HOP. Routes learned from one iBGP peer are not re-advertised to other iBGP peers (the iBGP split-horizon rule), hence full mesh or route reflectors.
Peer / neighborConfigured explicitly by IP and remote AS. BGP never discovers neighbours; a typo in the AS number keeps the session in Active forever.
RIB-in, Loc-RIB, RIB-outRoutes received per peer (after inbound policy), the chosen best paths, and what is sent per peer (after outbound policy).

The session: messages and the finite state machine

Idle ──start──▶ Connect ──TCP up──▶ OpenSent ──OPEN rcvd, OK──▶ OpenConfirm ──KEEPALIVE──▶ Established ▲ │ TCP fails │ └── NOTIFICATION ◀┴── Active (retrying TCP; "stuck in Active" = can't reach peer / wrong AS)◀┘
MessagePurposeWhat you'd check
OPENMy AS, hold time, router ID, capabilities (MP-BGP address families, 4-byte AS, route refresh, graceful restart)AS mismatch or capability mismatch → NOTIFICATION and back to Idle
KEEPALIVESent every hold/3 (default 60 s, hold 180 s). No keepalive within hold time → session torn down, all routes from that peer withdrawnFlapping sessions → look for packet loss or CPU starvation, not BGP config
UPDATEAdvertise prefixes with their attributes, or withdraw prefixesIncremental: after the initial table dump only changes are sent
NOTIFICATIONError, then close the session. Codes tell you why (e.g. "hold timer expired", "bad peer AS", "cease / max prefix")Read the last notification in show bgp neighbor before anything else

Before any of this: the TCP session must come up. "BGP is down" troubleshooting starts with: can I ping the peer address? Is port 179 open (ss -tn | grep :179, telnet peer 179)? Is the eBGP peer directly connected or do I need ebgp-multihop? Is the update source the address the peer expects? Is there an ACL or MD5/TCP-AO mismatch?

Path attributes (the ones that decide the best path)

AttributeTypeSet byCarriedWhat it means / how operators use it
WEIGHTCisco-only, local to one routerlocal policynot at allHighest wins. Quick per-router override.
LOCAL_PREFwell-known discretionaryinbound policy at the AS edgewithin the AS (iBGP)Highest wins. How an AS picks its exit (outbound traffic). Default 100.
AS_PATHwell-known mandatoryeach AS prepends itself on eBGP egresseverywhereShorter wins. Loop prevention: drop any route whose AS_PATH already contains my AS. Prepending makes a path look longer to discourage inbound traffic.
ORIGINwell-known mandatorythe originatoreverywhereIGP (i) < EGP (e) < incomplete (?). Lower wins; rarely decisive today.
MED (multi-exit discriminator)optional non-transitivethe advertising ASto the neighbouring AS onlyLowest wins. "If you have several links into me, use this one." Only compared between paths from the same neighbouring AS by default. A hint for inbound traffic; the neighbour may ignore it.
NEXT_HOPwell-known mandatoryeBGP: the advertising router; iBGP: unchangedeverywhereMust be reachable via the IGP or a static route, or the route is unusable. Classic iBGP bug: fix with next-hop-self on the border router.
COMMUNITYoptional transitiveanyoneas far as policies pass itA 32-bit tag (AS:value) used to signal policy: "set local-pref 80", "do not export", "prepend twice toward peers". NO_EXPORT (keep inside the AS) and NO_ADVERTISE are well-known. Large communities are the 96-bit version.

Best-path selection (say it in order)

  1. Ignore the route if the NEXT_HOP is unreachable or it fails a policy (e.g. AS_PATH contains my AS).
  2. Highest WEIGHT (Cisco).
  3. Highest LOCAL_PREF.
  4. Locally originated routes (network / aggregate / redistribute) before learned ones.
  5. Shortest AS_PATH.
  6. Lowest ORIGIN (i < e < ?).
  7. Lowest MED (only among paths from the same neighbour AS, unless always-compare-med).
  8. eBGP over iBGP.
  9. Lowest IGP metric to the NEXT_HOP ("hot-potato": leave the AS at the nearest exit).
  10. If multipath is enabled and the paths tie here → ECMP (install several). Otherwise:
  11. Oldest route (for eBGP, to reduce flapping), then lowest router ID, then lowest neighbour address.

A mnemonic many people use: "We Love Oranges AS Oranges Mean Pure Refreshment" (Weight, Local_pref, Originated, AS_path, Origin, MED, Paths external, RID). The exact tail differs by vendor; say "implementation-specific after the IGP metric step" and nobody will argue.

Steering traffic: the four knobs

GoalKnobWhy it works
Choose how traffic leaves my AS (outbound)LOCAL_PREF on inbound policy from each providerIt is compared early and shared across all my routers via iBGP, so the whole AS agrees on the exit.
Influence how traffic enters my AS (inbound)AS_PATH prepending on outbound advertisements to the less-preferred provider; MED toward a single provider with several links; communities the provider honours; or advertise more specific prefixes on the preferred linkI cannot set other ASes' LOCAL_PREF, so I can only make my route look worse elsewhere. More-specifics win on longest-prefix match regardless of BGP attributes, which is why they are the strongest (and most abused) tool.
Filter what I acceptprefix lists (no RFC1918, no default, no too-long prefixes), max-prefix limit, AS_PATH filters, RPKI validationProtects you from route leaks and hijacks.
Filter what I announceonly my own and customers' prefixes; NO_EXPORT on internal onesAnnouncing a provider's routes to another provider makes you a transit AS (a route leak) and can black-hole their traffic through your links.

iBGP scaling: route reflectors and confederations

iBGP won't re-advertise iBGP-learned routes, so every border router must peer with every other: n(n−1)/2 sessions. A route reflector (RR) is allowed to re-advertise iBGP routes to its clients; clients peer only with the RRs (usually two for redundancy). Loop prevention replaces the split-horizon rule with the ORIGINATOR_ID and CLUSTER_LIST attributes. Trade-off: an RR picks its best path and reflects only that, so clients may see a less-optimal exit (fixable with add-path or placing RRs well). Confederations split a big AS into sub-ASes that run eBGP-like sessions between them while appearing as one AS outside. Reflectors are far more common.

Data-center fabrics avoid the problem entirely: every leaf-spine link is an eBGP session with its own private ASN, no iBGP at all. See topic 16.

Convergence, flapping and safety

What you'd type

! Cisco IOS-style sketch
router bgp 65001
 bgp router-id 10.0.0.1
 neighbor 203.0.113.2 remote-as 65002           ! eBGP, directly connected
 neighbor 10.0.0.2 remote-as 65001              ! iBGP peer
 neighbor 10.0.0.2 update-source Loopback0
 neighbor 10.0.0.2 next-hop-self
 network 198.51.100.0 mask 255.255.255.0        ! must exist in the IGP/RIB to be advertised
 neighbor 203.0.113.2 route-map FROM-ISP in
 neighbor 203.0.113.2 maximum-prefix 1000000 90

show bgp summary            ! peer state, prefixes received; "Active"/"Idle" = not established
show bgp ipv4 unicast 198.51.100.0/24   ! all paths, the chosen one marked >, attributes
show bgp neighbors 203.0.113.2 advertised-routes / received-routes

# FRRouting on Linux (what a Meta-style fabric switch runs):
vtysh -c 'show bgp summary'
vtysh -c 'show bgp ipv4 unicast 10.2.2.0/24'

Interview questions

1. What is BGP and why does the Internet use it instead of OSPF?BGP is the path-vector protocol that exchanges reachability between autonomous systems over TCP. OSPF would need every router to hold the whole topology and trust every other router; it has no notion of policy or administrative boundaries and would never scale to a million prefixes and tens of thousands of untrusting organisations. BGP hides topology (only the AS path is visible), applies policy per neighbour and scales because it is incremental and TCP-based.
2. eBGP vs iBGP: what differs?eBGP peers are in different ASes, usually directly connected with TTL 1, prepend their AS and rewrite the next hop. iBGP peers are in the same AS, often loopback-to-loopback across the IGP, do not change AS_PATH or NEXT_HOP, and do not re-advertise routes learned from other iBGP peers, so you need a full mesh or route reflectors. Administrative distance differs (20 vs 200 on Cisco). The best-path rule prefers eBGP over iBGP.
3. How does BGP prevent loops?eBGP: a router discards any UPDATE whose AS_PATH already contains its own AS. iBGP: routes learned from an iBGP peer are not re-advertised to iBGP peers; route reflectors add ORIGINATOR_ID and CLUSTER_LIST to detect loops when they reflect.
4. Walk through session establishment.Both sides configure each other by IP and AS. TCP connects to port 179 (the lower router ID becomes the passive side if both connect). OPEN messages exchange AS, hold time, router ID and capabilities; if they agree, KEEPALIVEs confirm and the state is Established. Then each side sends its full table in UPDATEs, and after that only changes. The states are Idle, Connect, Active, OpenSent, OpenConfirm, Established.
5. A peer is stuck in Active. What do you check?Active means TCP can't be completed. Check IP reachability to the peer (ping), whether 179 is blocked (ACL, firewall), whether the eBGP peer is more than one hop away without multihop, whether the update source matches what the peer expects, and whether the peer even has me configured. A mismatched AS shows briefly as OpenSent then a NOTIFICATION, not Active.
6. List the best-path attributes in order.Weight, local preference, locally originated, AS path length, origin, MED, eBGP over iBGP, IGP metric to next hop, then tie-breakers (oldest, router ID, neighbour address). Say that the next hop must be reachable before any of this applies.
7. How do you control outbound vs inbound traffic?Outbound is easy: set LOCAL_PREF inbound from the preferred provider and iBGP spreads it. Inbound is a suggestion: prepend AS_PATH toward the less-preferred provider, send MED to a provider with multiple links, use the provider's communities, or advertise more-specific prefixes on the preferred link. Nothing forces another AS; its policy wins on its side.
8. What is MED and why is it weak?Multi-exit discriminator: a metric sent to a neighbouring AS to say which of several links into me to prefer. It is compared only after local-pref and AS-path, only between paths from the same neighbour AS, and the neighbour may reset or ignore it.
9. What is NEXT_HOP and the classic iBGP next-hop problem?The address traffic should be sent to for that prefix. eBGP sets it to the advertising router. iBGP leaves it unchanged, so internal routers see an external address they may have no route to and the route is invalid. Fix with next-hop-self on the border router or by carrying the external link in the IGP.
10. What is a route reflector and what does it cost you?An iBGP router allowed to re-advertise iBGP routes to its clients, removing the full-mesh requirement. It reflects only its own best path, so clients can lose path diversity and pick suboptimal exits; add-path or careful placement mitigates. It is also a single point of failure, so deploy in pairs.
11. What is a route leak and how do you prevent it?Advertising routes beyond their intended scope, e.g. passing one provider's routes to another so you become transit for them. Prevent with outbound prefix lists (only your own and customers'), communities that tag route origin, max-prefix limits on peers, and RPKI/ROV so invalid origins are dropped.
12. Why does Meta run BGP inside its data centers rather than an IGP?Scale and operational control: thousands of switches would strain a single link-state database, and BGP gives per-link policy, simple eBGP between tiers with private ASNs, ECMP across spines, and a small, well-understood implementation they could write themselves. The 2021 Meta paper on this is the thing to cite. See topic 16.
13. Compare BGP with OSPF in one minute.OSPF: link-state IGP, floods LSAs, every router computes SPF over the full topology, fast convergence, metric is cost, trusts all routers, runs directly over IP protocol 89, areas limit the database. BGP: path-vector EGP, advertises prefixes with attributes, policy-driven best path, slower convergence, runs over TCP 179, scales to the Internet, hides topology. Inside a small network use OSPF; between organisations or at hyperscale use BGP.

Traps

Scenario

Your AS has two providers. After a maintenance, all outbound traffic uses provider B, but you want A. Internally, routers show the A-learned routes but not as best. How do you investigate and fix?

Expected reasoning
  1. show bgp ipv4 unicast <prefix> on a border router: compare the attributes of the A path and the B path. Look at LOCAL_PREF first; if B's inbound route-map sets 200 and A's sets 100 (or the maintenance dropped A's route-map), that explains it.
  2. If LOCAL_PREF ties, check AS_PATH length (did A start prepending?), then whether A's NEXT_HOP is reachable (a missing next-hop-self or IGP route makes A's paths invalid, which looks like "there but not best").
  3. Fix by re-applying the inbound policy on the A session (set local-preference 200), then clear bgp neighbor soft in so it re-evaluates without resetting the session.
  4. Verify: best-path marker moves to A on every router (iBGP carries local-pref), traffic shifts (interface counters), and no prefixes are now unreachable.

← 13 · Graphs: BFS and DFS on grids and adjacency lists · all topics · 15 · JSON, CSV, scripts and the automation mindset →