14 · BGP in depth Network
AS, eBGP vs iBGP, TCP 179, FSM, UPDATE attributes, best-path order, policy, route reflectors, loop prevention
Why it matters for NPE. BGP is the most-reported networking topic for Meta NPE. Pick it as your 'protocol I know best' and be ready for 15 minutes of follow-ups.
Primer: BGP is a policy protocol that happens to find paths
Inside one organisation an IGP (OSPF, IS-IS) finds the shortest path and trusts every router. Between organisations nobody trusts anybody and "shortest" is not the goal: money, contracts and traffic engineering are. BGP (Border Gateway Protocol, RFC 4271) is a path-vector protocol: each route carries the list of autonomous systems it has crossed, and every router applies policy to decide what to accept, what to prefer and what to tell its neighbours. It runs over TCP port 179, so it inherits reliability and needs no timers for retransmission of its own.
Watch
First pass: eBGP vs iBGP, how neighbours form (TCP 179, the states), the attribute list and best-path order, a basic eBGP config. Watch all of it, then read the tables below with the video paused.
Depth. Watch 19:46-36:16 (basics, confederations, route reflectors, neighbour formation) and 42:24-1:00:03 (path selection). Skip the synchronisation and lab demo sections.
Weight, local preference, AS-path prepending and MED with labs. Say "local preference", not "weight", when asked for a vendor-neutral knob.
Meta's own talk: why eBGP-only in the fabric, ASN per tier, summarisation, a tiny policy set, an in-house BGP agent tested like software. Watch whole; it is your "why Meta runs BGP in the DC" answer.
Optional operator-level depth: policy with local-pref, MED and communities, route reflectors, when to use BGP instead of an IGP. The first hour is the core.
Reading that beats video here: RFC 4271 §9.1 (the decision process), RFC 7938 §5-6 (BGP in large data centers) and Meta's 2021 post on BGP in the data center.
Vocabulary
| Term | Meaning |
|---|---|
| AS (autonomous system) | A network under one administration with one routing policy, identified by an ASN (16-bit, now 32-bit; 64512-65534 and 4200000000+ are private). Meta is AS32934. |
| Prefix / NLRI | What BGP advertises: a network with its length (203.0.113.0/24) plus path attributes. NLRI = Network Layer Reachability Information. |
| eBGP | A session between different ASes. Usually directly connected (TTL 1 by default). Prepends the local AS to AS_PATH and rewrites NEXT_HOP when advertising. |
| iBGP | A session inside one AS, carrying external routes across it. Does not change AS_PATH or NEXT_HOP. Routes learned from one iBGP peer are not re-advertised to other iBGP peers (the iBGP split-horizon rule), hence full mesh or route reflectors. |
| Peer / neighbor | Configured explicitly by IP and remote AS. BGP never discovers neighbours; a typo in the AS number keeps the session in Active forever. |
| RIB-in, Loc-RIB, RIB-out | Routes received per peer (after inbound policy), the chosen best paths, and what is sent per peer (after outbound policy). |
The session: messages and the finite state machine
| Message | Purpose | What you'd check |
|---|---|---|
| OPEN | My AS, hold time, router ID, capabilities (MP-BGP address families, 4-byte AS, route refresh, graceful restart) | AS mismatch or capability mismatch → NOTIFICATION and back to Idle |
| KEEPALIVE | Sent every hold/3 (default 60 s, hold 180 s). No keepalive within hold time → session torn down, all routes from that peer withdrawn | Flapping sessions → look for packet loss or CPU starvation, not BGP config |
| UPDATE | Advertise prefixes with their attributes, or withdraw prefixes | Incremental: after the initial table dump only changes are sent |
| NOTIFICATION | Error, then close the session. Codes tell you why (e.g. "hold timer expired", "bad peer AS", "cease / max prefix") | Read the last notification in show bgp neighbor before anything else |
Before any of this: the TCP session must come up. "BGP is down" troubleshooting starts with: can I ping the peer address? Is port 179 open (ss -tn | grep :179, telnet peer 179)? Is the eBGP peer directly connected or do I need ebgp-multihop? Is the update source the address the peer expects? Is there an ACL or MD5/TCP-AO mismatch?
Path attributes (the ones that decide the best path)
| Attribute | Type | Set by | Carried | What it means / how operators use it |
|---|---|---|---|---|
| WEIGHT | Cisco-only, local to one router | local policy | not at all | Highest wins. Quick per-router override. |
| LOCAL_PREF | well-known discretionary | inbound policy at the AS edge | within the AS (iBGP) | Highest wins. How an AS picks its exit (outbound traffic). Default 100. |
| AS_PATH | well-known mandatory | each AS prepends itself on eBGP egress | everywhere | Shorter wins. Loop prevention: drop any route whose AS_PATH already contains my AS. Prepending makes a path look longer to discourage inbound traffic. |
| ORIGIN | well-known mandatory | the originator | everywhere | IGP (i) < EGP (e) < incomplete (?). Lower wins; rarely decisive today. |
| MED (multi-exit discriminator) | optional non-transitive | the advertising AS | to the neighbouring AS only | Lowest wins. "If you have several links into me, use this one." Only compared between paths from the same neighbouring AS by default. A hint for inbound traffic; the neighbour may ignore it. |
| NEXT_HOP | well-known mandatory | eBGP: the advertising router; iBGP: unchanged | everywhere | Must be reachable via the IGP or a static route, or the route is unusable. Classic iBGP bug: fix with next-hop-self on the border router. |
| COMMUNITY | optional transitive | anyone | as far as policies pass it | A 32-bit tag (AS:value) used to signal policy: "set local-pref 80", "do not export", "prepend twice toward peers". NO_EXPORT (keep inside the AS) and NO_ADVERTISE are well-known. Large communities are the 96-bit version. |
Best-path selection (say it in order)
- Ignore the route if the NEXT_HOP is unreachable or it fails a policy (e.g. AS_PATH contains my AS).
- Highest WEIGHT (Cisco).
- Highest LOCAL_PREF.
- Locally originated routes (network / aggregate / redistribute) before learned ones.
- Shortest AS_PATH.
- Lowest ORIGIN (i < e < ?).
- Lowest MED (only among paths from the same neighbour AS, unless always-compare-med).
- eBGP over iBGP.
- Lowest IGP metric to the NEXT_HOP ("hot-potato": leave the AS at the nearest exit).
- If multipath is enabled and the paths tie here → ECMP (install several). Otherwise:
- Oldest route (for eBGP, to reduce flapping), then lowest router ID, then lowest neighbour address.
A mnemonic many people use: "We Love Oranges AS Oranges Mean Pure Refreshment" (Weight, Local_pref, Originated, AS_path, Origin, MED, Paths external, RID). The exact tail differs by vendor; say "implementation-specific after the IGP metric step" and nobody will argue.
Steering traffic: the four knobs
| Goal | Knob | Why it works |
|---|---|---|
| Choose how traffic leaves my AS (outbound) | LOCAL_PREF on inbound policy from each provider | It is compared early and shared across all my routers via iBGP, so the whole AS agrees on the exit. |
| Influence how traffic enters my AS (inbound) | AS_PATH prepending on outbound advertisements to the less-preferred provider; MED toward a single provider with several links; communities the provider honours; or advertise more specific prefixes on the preferred link | I cannot set other ASes' LOCAL_PREF, so I can only make my route look worse elsewhere. More-specifics win on longest-prefix match regardless of BGP attributes, which is why they are the strongest (and most abused) tool. |
| Filter what I accept | prefix lists (no RFC1918, no default, no too-long prefixes), max-prefix limit, AS_PATH filters, RPKI validation | Protects you from route leaks and hijacks. |
| Filter what I announce | only my own and customers' prefixes; NO_EXPORT on internal ones | Announcing a provider's routes to another provider makes you a transit AS (a route leak) and can black-hole their traffic through your links. |
iBGP scaling: route reflectors and confederations
iBGP won't re-advertise iBGP-learned routes, so every border router must peer with every other: n(n−1)/2 sessions. A route reflector (RR) is allowed to re-advertise iBGP routes to its clients; clients peer only with the RRs (usually two for redundancy). Loop prevention replaces the split-horizon rule with the ORIGINATOR_ID and CLUSTER_LIST attributes. Trade-off: an RR picks its best path and reflects only that, so clients may see a less-optimal exit (fixable with add-path or placing RRs well). Confederations split a big AS into sub-ASes that run eBGP-like sessions between them while appearing as one AS outside. Reflectors are far more common.
Data-center fabrics avoid the problem entirely: every leaf-spine link is an eBGP session with its own private ASN, no iBGP at all. See topic 16.
Convergence, flapping and safety
- Session loss means all routes from that peer are withdrawn, so a 180 s hold time is a long outage. Operators use BFD (sub-second failure detection) or short timers on direct links.
- Route flap dampening penalises prefixes that keep flapping and suppresses them for a while. Mostly disabled today because it punishes victims of a single flapping link upstream.
- Graceful restart keeps forwarding while the BGP process restarts. Max-prefix tears down a session that suddenly sends far more routes than expected (a full table leaked by mistake).
- The full Internet table is around a million IPv4 prefixes; a default-free router needs the FIB capacity and the memory for it.
- Meta's October 2021 outage: a maintenance command withdrew the backbone routes, the data centers lost connectivity, and the DNS servers, by design, withdrew their own BGP advertisements when they could not reach the data centers, so the outside world saw DNS failure while the cause was routing. Good case study for "symptom vs cause".
What you'd type
! Cisco IOS-style sketch
router bgp 65001
bgp router-id 10.0.0.1
neighbor 203.0.113.2 remote-as 65002 ! eBGP, directly connected
neighbor 10.0.0.2 remote-as 65001 ! iBGP peer
neighbor 10.0.0.2 update-source Loopback0
neighbor 10.0.0.2 next-hop-self
network 198.51.100.0 mask 255.255.255.0 ! must exist in the IGP/RIB to be advertised
neighbor 203.0.113.2 route-map FROM-ISP in
neighbor 203.0.113.2 maximum-prefix 1000000 90
show bgp summary ! peer state, prefixes received; "Active"/"Idle" = not established
show bgp ipv4 unicast 198.51.100.0/24 ! all paths, the chosen one marked >, attributes
show bgp neighbors 203.0.113.2 advertised-routes / received-routes
# FRRouting on Linux (what a Meta-style fabric switch runs):
vtysh -c 'show bgp summary'
vtysh -c 'show bgp ipv4 unicast 10.2.2.0/24'
Interview questions
1. What is BGP and why does the Internet use it instead of OSPF?
BGP is the path-vector protocol that exchanges reachability between autonomous systems over TCP. OSPF would need every router to hold the whole topology and trust every other router; it has no notion of policy or administrative boundaries and would never scale to a million prefixes and tens of thousands of untrusting organisations. BGP hides topology (only the AS path is visible), applies policy per neighbour and scales because it is incremental and TCP-based.2. eBGP vs iBGP: what differs?
eBGP peers are in different ASes, usually directly connected with TTL 1, prepend their AS and rewrite the next hop. iBGP peers are in the same AS, often loopback-to-loopback across the IGP, do not change AS_PATH or NEXT_HOP, and do not re-advertise routes learned from other iBGP peers, so you need a full mesh or route reflectors. Administrative distance differs (20 vs 200 on Cisco). The best-path rule prefers eBGP over iBGP.3. How does BGP prevent loops?
eBGP: a router discards any UPDATE whose AS_PATH already contains its own AS. iBGP: routes learned from an iBGP peer are not re-advertised to iBGP peers; route reflectors add ORIGINATOR_ID and CLUSTER_LIST to detect loops when they reflect.4. Walk through session establishment.
Both sides configure each other by IP and AS. TCP connects to port 179 (the lower router ID becomes the passive side if both connect). OPEN messages exchange AS, hold time, router ID and capabilities; if they agree, KEEPALIVEs confirm and the state is Established. Then each side sends its full table in UPDATEs, and after that only changes. The states are Idle, Connect, Active, OpenSent, OpenConfirm, Established.5. A peer is stuck in Active. What do you check?
Active means TCP can't be completed. Check IP reachability to the peer (ping), whether 179 is blocked (ACL, firewall), whether the eBGP peer is more than one hop away without multihop, whether the update source matches what the peer expects, and whether the peer even has me configured. A mismatched AS shows briefly as OpenSent then a NOTIFICATION, not Active.6. List the best-path attributes in order.
Weight, local preference, locally originated, AS path length, origin, MED, eBGP over iBGP, IGP metric to next hop, then tie-breakers (oldest, router ID, neighbour address). Say that the next hop must be reachable before any of this applies.7. How do you control outbound vs inbound traffic?
Outbound is easy: set LOCAL_PREF inbound from the preferred provider and iBGP spreads it. Inbound is a suggestion: prepend AS_PATH toward the less-preferred provider, send MED to a provider with multiple links, use the provider's communities, or advertise more-specific prefixes on the preferred link. Nothing forces another AS; its policy wins on its side.8. What is MED and why is it weak?
Multi-exit discriminator: a metric sent to a neighbouring AS to say which of several links into me to prefer. It is compared only after local-pref and AS-path, only between paths from the same neighbour AS, and the neighbour may reset or ignore it.9. What is NEXT_HOP and the classic iBGP next-hop problem?
The address traffic should be sent to for that prefix. eBGP sets it to the advertising router. iBGP leaves it unchanged, so internal routers see an external address they may have no route to and the route is invalid. Fix with next-hop-self on the border router or by carrying the external link in the IGP.10. What is a route reflector and what does it cost you?
An iBGP router allowed to re-advertise iBGP routes to its clients, removing the full-mesh requirement. It reflects only its own best path, so clients can lose path diversity and pick suboptimal exits; add-path or careful placement mitigates. It is also a single point of failure, so deploy in pairs.11. What is a route leak and how do you prevent it?
Advertising routes beyond their intended scope, e.g. passing one provider's routes to another so you become transit for them. Prevent with outbound prefix lists (only your own and customers'), communities that tag route origin, max-prefix limits on peers, and RPKI/ROV so invalid origins are dropped.12. Why does Meta run BGP inside its data centers rather than an IGP?
Scale and operational control: thousands of switches would strain a single link-state database, and BGP gives per-link policy, simple eBGP between tiers with private ASNs, ECMP across spines, and a small, well-understood implementation they could write themselves. The 2021 Meta paper on this is the thing to cite. See topic 16.13. Compare BGP with OSPF in one minute.
OSPF: link-state IGP, floods LSAs, every router computes SPF over the full topology, fast convergence, metric is cost, trusts all routers, runs directly over IP protocol 89, areas limit the database. BGP: path-vector EGP, advertises prefixes with attributes, policy-driven best path, slower convergence, runs over TCP 179, scales to the Internet, hides topology. Inside a small network use OSPF; between organisations or at hyperscale use BGP.Traps
- Saying BGP picks the fastest or lowest-latency path. It has no idea about latency; it picks by attributes and policy.
- Forgetting the NEXT_HOP reachability check: everything else is moot if the next hop isn't in the RIB.
- Claiming MED steers outbound traffic. MED is a hint to your neighbour about their traffic into you.
- Putting AS_PATH before LOCAL_PREF in the selection order.
- Saying iBGP peers must be directly connected. They usually peer over loopbacks across the IGP; eBGP is the one that defaults to TTL 1.
- Describing BGP as "layer 3". It is an application over TCP that installs layer-3 routes.
- Thinking a session reset is harmless: every route from that peer is withdrawn and traffic reroutes or drops.
Scenario
Your AS has two providers. After a maintenance, all outbound traffic uses provider B, but you want A. Internally, routers show the A-learned routes but not as best. How do you investigate and fix?
Expected reasoning
show bgp ipv4 unicast <prefix>on a border router: compare the attributes of the A path and the B path. Look at LOCAL_PREF first; if B's inbound route-map sets 200 and A's sets 100 (or the maintenance dropped A's route-map), that explains it.- If LOCAL_PREF ties, check AS_PATH length (did A start prepending?), then whether A's NEXT_HOP is reachable (a missing next-hop-self or IGP route makes A's paths invalid, which looks like "there but not best").
- Fix by re-applying the inbound policy on the A session (set local-preference 200), then
clear bgp neighbor soft inso it re-evaluates without resetting the session. - Verify: best-path marker moves to A on every router (iBGP carries local-pref), traffic shifts (interface counters), and no prefixes are now unreachable.
← 13 · Graphs: BFS and DFS on grids and adjacency lists · all topics · 15 · JSON, CSV, scripts and the automation mindset →