16 · Meta's network: Clos fabrics, ECMP, BGP in the DC, FBOSS, backbone and MPLS Network
Leaf/spine and oversubscription · ECMP hashing · why BGP inside a data center · F16, Minipack, FBOSS is Linux apps · Express Backbone, MPLS segment routing · RoCE for AI
Why it matters for NPE. Rarely asked directly, but it is the environment you'd work in. It answers 'what OS runs on a switch', powers a real 'Why Meta', and gives you questions for the interviewer. Keep it to one sitting.
Primer: the environment you would work in
Meta's network is three networks. The data center fabric is a Clos of commodity switches running Meta's own switch software, routed entirely with eBGP, with servers never more than a handful of hops apart. The backbone connects data centers to each other (Express Backbone) and to the Internet (Classic Backbone) with MPLS segment routing and centralised traffic engineering. The AI backend fabrics are separate Clos networks built for RDMA between GPUs. Nobody expects you to have operated any of it; they expect you to know the shape, the vocabulary, and why the design choices were made. That is enough for a real "Why Meta" and a few sharp questions back.
Watch
Alexey Andreyev presenting the 2014 fabric he designed: pods of 48 racks, four fabric switches per pod, four spine planes, edge pods, BGP as the only routing protocol, ECMP everywhere. The one Meta talk to know. Watch whole.
The F16 talk: 16 planes at 100G instead of 4 at 400G, Minipack as a single-ASIC 128x100G building block, HGRID, migrating a live fabric.
FBOSS in 2023: the switch agent, hardware abstraction via SAI, treating a switch like a server, continuous deployment and testing. Your "what OS runs on a switch" answer with evidence.
Already on the BGP page; rewatch after reading the "BGP in the DC" section below.
If leaf/spine is new: 14:17-18:36 for spine-leaf, after 4:33-14:17 on classic three-tier campus for contrast.
Labels, LER and LSR, label swapping, why a provider uses it. Enough for the MPLS section.
Optional: Express Backbone from the people who built it, traffic engineering included.
Optional: why Meta wrote its own routing platform. Useful if asked "why not just use a vendor".
Reading beats most of the video here. In order: Introducing data center fabric (2014), F16 and Minipack (2019), Running BGP in large-scale data centers (2021), FBOSS and Wedge in the open (2015), Building Express Backbone (2017), More details about the October 4 outage (2021), RoCE networks for distributed AI training (2024), and RFC 7938 §3 for Clos vocabulary.
Leaf/spine Clos
- Why: a classic three-tier campus (access, distribution, core) funnels traffic to a pair of big core boxes and relies on STP to block redundant links. A Clos fabric uses many small identical switches, every leaf has an equal-cost path through every spine, nothing is blocked, and every server is the same number of hops from every other, so latency is predictable and east-west traffic (server to server, which dominates in a data center) is not squeezed through a core.
- Non-blocking: if the sum of a leaf's uplink bandwidth equals the sum of its server-facing bandwidth, any traffic pattern can be served without contention. The oversubscription ratio is downlink bandwidth : uplink bandwidth at the leaf. 48 servers at 25G with 8 uplinks at 100G is 1200:800 = 3:2 (1.5:1). 1:1 is non-blocking; 3:1 or 4:1 is common where cost matters more.
- Horizontal scale: more bandwidth means more spines (more planes), not bigger boxes. More racks means more leaves. A five-stage Clos (pods of leaf/spine connected by a third tier) scales to a building.
- Failure domains: a spine failure costs you 1/N of the bandwidth to everyone, not an outage. A leaf failure costs you one rack. Design so that no single device is special; Meta's fabrics have no DR, no STP root, no stacked core.
ECMP
With N equal-cost next hops, the switch picks one per packet by hashing the 5-tuple (source IP, destination IP, protocol, source port, destination port) modulo N. The hash is per flow, not per packet: every packet of a TCP connection takes the same path, so no reordering, which TCP would read as loss. The price is imbalance: elephant flows (one big transfer) land on one link and can saturate it while neighbours idle, and polarisation happens when every tier uses the same hash function so the second tier sees only a subset of its inputs; fix with a per-switch hash seed. Resilient hashing (consistent hashing over the next-hop set) keeps existing flows on their paths when a member link fails or returns, instead of reshuffling every flow. BGP gives ECMP for free when paths tie through the IGP-metric step (topic 14), provided maximum-paths is set and the AS_PATH lengths match.
Why BGP inside the data center
RFC 7938 (Lapukhov et al., written from this experience) and Meta's 2021 post give the same answer; say it as five design points:
- eBGP only, no IGP, no iBGP. Every link between tiers is a single-hop eBGP session. No route reflectors, no LSDB, no flooding storms, no second protocol to keep consistent.
- Private ASN per tier or pod. Leaves in a pod share an ASN or each get one; each spine plane has its own; the AS_PATH itself prevents loops and tells you which tier a route came from.
allowas-inoras-overrideis avoided by design. - /31 point-to-point links (or unnumbered / IPv6 link-local, RFC 5549) so a pod of hundreds of links fits in a small address block and addressing is generated, not planned.
- Hierarchical summarisation. Each rack announces its prefix, each pod summarises its racks, each plane summarises pods. FIBs stay small enough for merchant-silicon ASICs, and a flapping rack does not churn the whole fabric. The 2021 post pairs this with policy for backup paths so a summary does not black-hole a withdrawn more-specific.
- Tiny policy set and an in-house BGP agent. A handful of route-maps tagged with communities (tier, pod, backup) applied uniformly. Meta wrote its own BGP implementation (C++), with the features it needs and nothing else, and tests, canaries and deploys it like any other service. Convergence relies on BFD plus fast external fallover, not hold timers.
The honest counterpoint an interviewer may push on: an IGP would converge faster out of the box and needs less configuration per link. Meta's answer is that configuration is generated by software anyway, and that per-link sessions with explicit policy are easier to reason about and to debug at tens of thousands of switches. Quote the 2021 post: BGP works in the DC only with "tight codesign with the data center topology, configuration, switch software, and operational pipeline".
Meta's published architecture, at the level you should know
| Piece | Year | What to say |
|---|---|---|
| Data center fabric | 2014 | Replaced clusters built around giant chassis switches. Server pods of 48 racks, each rack switch uplinked to 4 fabric switches; 4 independent spine planes connect every pod's fabric switches; edge pods connect to the backbone. Standard BGP4 as the only routing protocol, ECMP everywhere, all topology and config generated by automation. |
| F16 and Minipack | 2019 | Scaled the same idea: 16 spine planes at 100G instead of 4 at 400G, so each rack gets 1.6 Tbit/s uplink with 100G optics that existed and 400G-ready cabling. Minipack is the single building block for every tier: one 12.8T ASIC, 128 x 100G modular ports, 4 RU, replacing a multi-chip chassis (Backpack). Later Minipack2 (25.6T, 200G ports) and Minipack3 (51.2T). HGRID is the aggregation that connects fabrics across a region. |
| FBOSS | 2015 onward | "FBOSS is not a full operating system. Rather, it is a set of applications that can be run on a standard Linux OS." The switch boots Linux (CentOS-derived) on an x86 CPU; the FBOSS agent programs the ASIC through SAI (the Switch Abstraction Interface), speaks Thrift to the rest of Meta's tooling, and is deployed and monitored exactly like a server service. OpenBMC runs the board management controller. A vendor NOS (NX-OS, Junos, EOS) is the same picture with the vendor's apps on top of their Linux or BSD. That is the full answer to "what OS runs on a switch". |
| Express Backbone (EBB) | 2017 | Private DC-to-DC backbone split from the Internet-facing Classic Backbone. Open/R (Meta's in-house link-state routing platform) as the IGP, MPLS segment routing labels programmed by LSP agents on each router, and a centralised traffic-engineering controller that computes paths from demand and capacity and pushes them down. Hybrid: distributed protocol for liveness, central brain for optimisation. |
| October 2021 outage | 2021 | A routine backbone maintenance command, wrongly passed by an audit tool, withdrew all backbone routes and disconnected every data center from each other and the Internet. Meta's authoritative DNS servers are designed to withdraw their own BGP announcements when they cannot reach the data centers, so the public symptom was DNS failure while the cause was routing, and the fix required people physically at the sites. Lesson: symptom versus cause, blast radius of automation, and keep out-of-band access that does not depend on the production network. |
| AI backend fabrics | 2024 | Separate frontend (normal traffic) and backend (GPU-to-GPU) networks. The backend is RoCEv2 (RDMA over Converged Ethernet, RDMA inside UDP) on a two-stage Clos of rack switches (RTSW), cluster switches (CTSW) and aggregators (ATSW). Plain ECMP balanced badly because training has few, huge flows, so they moved to path pinning and then enhanced ECMP with more queue pairs; lossless delivery via PFC and receiver-driven admission control rather than DCQCN at 400G. |
MPLS basics
MPLS forwards on a 20-bit label in a 4-byte shim header between the layer-2 header and the IP header (label, 3 bits traffic class, 1 bit bottom-of-stack, 8 bits TTL). The ingress router (LER) classifies the packet once and pushes a label; each core router (LSR) looks up only the label, swaps it and forwards; the egress pops it, or the second-to-last router pops it (penultimate hop popping, PHP) so the egress does one lookup instead of two. The sequence of label switches a packet follows is a label-switched path (LSP). Labels used to be distributed by LDP or RSVP-TE. Three reasons it exists: traffic engineering (steer a flow down a path the IGP would not choose), VPNs (stack a second label that identifies the customer VRF, so one backbone carries many isolated networks: L3VPN, EVPN), and fast reroute (pre-computed backup LSPs switch in under 50 ms). Segment routing (SR-MPLS) keeps the labels but drops LDP: the IGP (IS-IS or OSPF, or Open/R at Meta) advertises a label per node and per link, and the ingress pushes a stack of labels that spells out the path, so the core keeps no per-path state. That is what EBB programs.
Ethernet | MPLS label 16005 (TC 0, S=1, TTL 63) | IP 10.1.1.1 → 10.2.2.2 | TCP ...
ingress pushed core swaps 16005 → 16012 egress (or PHP) pops
Why Meta, and questions to ask the interviewer
The honest version of "Why Meta" for this role: the network is built from software, published in the open, and operated by the people who write it. An intern here touches BGP, Linux, Python and production incidents on the same day. Questions that show you read the material and want to know how the job actually feels:
- How much of an NPE's week is writing tooling versus responding to the fabric or backbone, and does that split change between the data center and backbone teams?
- With FBOSS deployed like a service, what does a bad switch-software rollout look like from the operator's seat, and how far does the canary go before it reaches a whole plane?
- After 2021, what changed in how maintenance commands on the backbone are audited and in what out-of-band access exists when the production network is gone?
- The RoCE post says plain ECMP was not good enough for training traffic. Are the AI backend fabrics operated by the same NPE teams as the frontend fabrics, and what is different about on-call for them?
- What does a new engineer break first, and what tooling exists so an intern can safely make a change to production?
Interview questions
1. What operating system runs on a switch or router?
Almost always Linux or BSD underneath, with the vendor's network applications on top and a driver layer that programs the forwarding ASIC. Meta's FBOSS is explicit about it: a standard Linux distribution with the FBOSS agent as one more application, programming the ASIC through SAI, deployed and monitored like a server. Cisco NX-OS, Arista EOS and Juniper Junos (FreeBSD, now Linux in Junos Evolved) are the same shape with vendor apps.2. Describe a leaf/spine fabric and why data centers use it.
Two tiers: every leaf (top of rack) connects to every spine, leaves never to leaves, spines never to spines. Every server-to-server path is leaf, spine, leaf, so hop count and latency are uniform, all links are active with ECMP (no STP blocking), bandwidth scales by adding spines, and a single spine failure costs a fraction of capacity instead of an outage. It fits east-west heavy traffic better than a three-tier campus funnelled through a core pair.3. What is oversubscription?
The ratio of server-facing bandwidth to uplink bandwidth at a leaf. 48 x 25G down and 8 x 100G up is 1.5:1. 1:1 is non-blocking; higher ratios save money on the assumption that not every server bursts at once. Follow-up: where would you accept 4:1? Storage or web tiers with bursty, mostly-local traffic; never on a training cluster.4. How does ECMP pick a path, and what are its failure modes?
Hash the 5-tuple modulo the number of equal-cost next hops, so each flow sticks to one path and TCP sees no reordering. Failure modes: elephant flows overload one link while others idle; polarisation when every tier hashes identically; and a member failing reshuffles all flows unless resilient hashing is used. Fixes: per-switch hash seeds, resilient or consistent hashing, flowlet switching, or at the application level more flows.5. Why does Meta run BGP in the data center instead of OSPF?
Scale and operability: tens of thousands of switches would make one link-state database and its flooding unmanageable, while eBGP gives one simple session per link, loop prevention from the AS_PATH, summarisation per pod and plane to keep FIBs small, a tiny uniform policy set, and ECMP. Meta also wanted an implementation small enough to write and test themselves. RFC 7938 and the 2021 post are the references; see topic 14 for the protocol.6. What is F16 in two sentences?
Meta's 2019 fabric design: 16 spine planes at 100G giving each rack 1.6 Tbit/s of uplink with optics that existed, instead of 4 planes at 400G. Every tier is built from the same box, Minipack, a single-ASIC 128 x 100G switch running FBOSS.7. What is MPLS and why would a backbone use it?
Forwarding on a 20-bit label pushed at ingress, swapped in the core and popped at or before egress, along a label-switched path. It lets you steer traffic along paths the IGP would not pick (traffic engineering), carry many isolated networks over one core with stacked labels (VPNs), and pre-compute backup paths for sub-50 ms reroute. Segment routing keeps the labels but removes LDP by having the IGP advertise them, and the ingress encodes the path as a label stack.8. What happened in the October 2021 outage and what does it teach?
A backbone maintenance command withdrew all backbone routes, cutting every data center off; the DNS servers, by design, withdrew their BGP announcements when they lost the data centers, so the world saw DNS failure while the cause was routing. Lessons: symptom is not cause, automation needs guardrails with real blast-radius limits, and out-of-band access must not depend on the network it manages.9. What is RoCE and why does an AI cluster need a separate fabric?
RDMA over Converged Ethernet: GPUs write directly into each other's memory across the network, carried in UDP (RoCEv2), and it assumes a near-lossless network. Training traffic is a few enormous synchronised flows, which ECMP balances badly and which cannot tolerate drops, so Meta builds a dedicated backend Clos with PFC, receiver-driven admission control and tuned load balancing, separate from the frontend fabric that carries ordinary traffic.Traps
- Saying ECMP is per packet. It is per flow; per-packet spraying reorders TCP and looks like loss.
- Saying leaves connect to each other for redundancy. They do not; redundancy comes from multiple spines.
- Calling FBOSS an operating system. It is a set of applications on Linux; the quote is in the 2015 post.
- Saying Meta runs iBGP or OSPF inside the fabric. eBGP only, private ASNs, one session per link.
- Describing the 2021 outage as "a DNS outage". DNS was the symptom; a backbone route withdrawal was the cause.
- Claiming MPLS encrypts or that it is a VPN by itself. Labels isolate forwarding; privacy is separation, not encryption.
- Saying segment routing needs LDP. The whole point is that the IGP carries the labels and the ingress holds the path state.
Scenario
In a leaf/spine fabric with 16 spine planes, one rack reports intermittent packet loss to a specific set of destination racks, roughly 1 in 16 of its flows. Other racks are fine. What do you suspect and how do you confirm it?
Expected reasoning
- "About 1 in 16 flows, specific destinations" fits one bad path out of 16 ECMP next hops: a single plane's link from this leaf, or one spine in that plane, is dropping. Per-flow hashing pins unlucky flows to it, which is why the loss is intermittent per flow and steady per path.
- On the leaf: check interface counters and optics on all 16 uplinks (errors, CRCs, discards, light levels); check the BGP sessions (
show bgp summary) and whether all 16 next hops are installed for the affected prefixes. - Reproduce by varying source ports toward an affected destination (
mtrortraceroutewith many source ports, or a small script) and find the port set that fails; the path they share names the plane. - Fix: drain the suspect link or spine (shut the eBGP session, or lower its preference via the fabric tooling) so ECMP excludes it, confirm loss clears, then repair the optic or cable and un-drain.
- Say what the fabric design bought you: a drained plane costs 1/16 of bandwidth, nobody else noticed, and the fix was a routing-policy action, not a maintenance window.
← 15 · JSON, CSV, scripts and the automation mindset · all topics · 17 · Classics: binary search, recursion, trees and BST →