10 · Switching: MAC learning, VLANs, STP and LAG Network
Flood/learn/forward · 802.1Q trunks · root bridge and blocked ports · LACP and hashing
Why it matters for NPE. Layer 2 questions (MAC tables, VLAN mismatches, loops) appear in Menlo Park reports and are where many software-leaning candidates are weakest.
Primer: a switch does three things to a frame
A switch learns the source MAC of every frame and the port it arrived on, forwards a frame whose destination MAC it knows out that one port, and floods everything else (unknown unicast, broadcast, most multicast) out every other port in the same VLAN. That is the whole data plane. VLANs split one switch into several of these broadcast domains. STP exists because flooding plus a physical loop is fatal. LAG exists because you want two links between switches without STP blocking one of them.
Watch
Learn, flood, forward and filter, shown frame by frame as the MAC table fills. Eleven minutes; watch the whole thing and redraw the table on paper afterwards.
Watch 0:00-13:05: 2:02 broadcast domains, 4:38 what a VLAN is, 6:22 segmenting at layer 3, 10:07 segmenting at layer 2. The rest is Cisco CLI.
Watch 4:26-16:53: 4:26 trunk ports, 8:15 the 802.1Q tag, 9:12 TPID, 10:14 PCP, 11:02 VID, 12:22 VLAN ranges, 13:44 native VLAN. 27:00 router-on-a-stick if inter-VLAN routing is new.
Watch 1:47-33:17: 4:15 loops and broadcast storms, 11:53 BPDUs and root election, 21:00 root port by cost, 24:21 by neighbor bridge ID, 26:48 by port ID, 29:13 blocking, 31:20 process summary.
Watch 1:48-17:57: 1:48 why bundle links, 8:59 load balancing by per-flow hash (why one flow never exceeds one member), 15:30 PAgP vs LACP vs static. Skip config from 17:57.
Only 2:19-7:21, the STP vs RSTP comparison table. The port role and state tables below cover the rest.
Optional refresher. 2:49 tagging and trunking, 4:40 native VLAN, 6:31 VXLAN as the bridge to data-center overlay vocabulary.
Learn, forward, flood: a worked MAC table
- Learning is on source address only. A switch never learns from destinations and never learns an IP address.
- Aging defaults to 300 s on most switches. A MAC that moves to another port is relearned instantly on the next frame from it; a MAC that seen on two ports alternately is flapping, which means a loop or a duplicated MAC.
- Table size is finite (thousands to hundreds of thousands of entries in CAM). Overflow turns the switch into a hub for new addresses, which is the point of MAC-flooding attacks and why port security limits MACs per port.
- Broadcast domain: the set of ports a broadcast reaches, one per VLAN, bounded by a router. Collision domain: the set of devices that share one medium; in a full-duplex switched network every port is its own collision domain and collisions only exist on hubs and half-duplex links. Say that to kill the question fast.
Switch# show mac address-table
Vlan Mac Address Type Ports
10 aaaa.aaaa.aaaa DYNAMIC Gi1/0/1
10 bbbb.bbbb.bbbb DYNAMIC Gi1/0/2
20 cccc.cccc.cccc DYNAMIC Gi1/0/3
10 0000.0c9f.f00a DYNAMIC Gi1/0/48 <-- gateway MAC (HSRP) learned on the uplink
$ bridge fdb show br0 | grep -v permanent # Linux bridge, same idea
aa:aa:aa:aa:aa:aa dev eth1 vlan 10 master br0
VLANs and 802.1Q
| Access port | Trunk port | |
|---|---|---|
| Carries | one VLAN, frames untagged on the wire | many VLANs, each frame tagged with its VID |
| Connects to | a host, server, printer, phone (voice VLAN is the one exception, tagged alongside the untagged data VLAN) | another switch, a router doing inter-VLAN routing, a hypervisor |
| Ingress | switch assigns the port's VLAN | switch reads the VID; untagged frames go to the native VLAN |
| Egress | tag removed | tag kept, except frames in the native VLAN which leave untagged |
| Config | switchport mode access, switchport access vlan 10 | switchport mode trunk, switchport trunk allowed vlan 10,20,30, switchport trunk native vlan 999 |
- Native VLAN is the one VLAN a trunk sends untagged, default 1. Both ends must agree or frames leak between VLANs: a frame leaving SW1 untagged in VLAN 1 arrives at SW2 and is placed in SW2's native VLAN, say 20. Good practice: set the native to an unused VLAN or tag it too (
vlan dot1q tag native). - Inter-VLAN routing. Hosts in VLAN 10 (10.1.10.0/24) and VLAN 20 (10.1.20.0/24) cannot talk without a router, because a VLAN is a broadcast domain and a broadcast domain is a subnet. Two ways: router-on-a-stick, one trunk to a router with a subinterface per VLAN (
eth0.10,eth0.20), every inter-VLAN packet goes up and back down the same link, so the link is the bottleneck. Or an SVI on a layer-3 switch (interface Vlan10with an IP), routed in the ASIC at line rate. Modern enterprise and all data-center designs use SVIs or routed ports. - The classic mismatch. A server port is set to VLAN 20 but the server's address is in VLAN 10's subnet. Link is up, the switch learns the MAC in VLAN 20, the server ARPs for a gateway that lives in VLAN 10, nobody answers, FAILED neighbor, or it gets a 169.254 address because the DHCP relay is on the VLAN 10 SVI. Second form: VLAN 30 is created on both switches but missing from
trunk allowed vlan, so hosts in VLAN 30 on SW1 can see each other and nothing on SW2.show interfaces trunkshows allowed and active VLANs per trunk; compare both ends.
$ tcpdump -eni eth0 vlan
aa:aa:aa:aa:aa:aa > ff:ff:ff:ff:ff:ff, ethertype 802.1Q (0x8100), length 46: vlan 10, p 0, ethertype ARP, Request who-has 10.1.10.1 tell 10.1.10.50
SW1# show interfaces trunk
Port Mode Encapsulation Status Native vlan
Gi1/0/48 on 802.1q trunking 999
Port Vlans allowed on trunk
Gi1/0/48 10,20 <-- VLAN 30 missing here is the whole bug
Spanning Tree: loops are fatal at layer 2
An Ethernet frame has no TTL. Connect two switches with two cables and send one broadcast: each switch floods it out the other link, the other switch floods it back, forever, at line rate. Within seconds the links are saturated (a broadcast storm), every switch sees the same source MAC arriving on two ports and rewrites its table on each frame (MAC flapping), unicast traffic is misdirected, and CPUs melt processing broadcasts. Hosts receive thousands of copies of every ARP. The only exit is pulling a cable. STP's job is to make the physical loop logically a tree by blocking ports.
| Port role | How chosen | State |
|---|---|---|
| Root port | One per non-root switch: the port with the lowest cost to the root. Ties: lowest neighbor bridge ID, then lowest neighbor port ID. | forwarding |
| Designated port | One per segment: the port on the switch with the lowest root cost on that segment (same tie-breaks). All root bridge ports are designated. | forwarding |
| Non-designated (802.1D) / Alternate, Backup (RSTP) | Everything else. Alternate = a second path to the root, Backup = a second port onto a segment I already serve. | blocking (RSTP: discarding). Still receives BPDUs so it can wake up. |
| Link speed | 802.1D short cost | RSTP/MST long cost |
|---|---|---|
| 10 Mbit/s | 100 | 2 000 000 |
| 100 Mbit/s | 19 | 200 000 |
| 1 Gbit/s | 4 | 20 000 |
| 10 Gbit/s | 2 | 2 000 |
| 100 Gbit/s | 1 (useless) | 200 |
- Classic 802.1D convergence: blocking → listening (15 s) → learning (15 s) → forwarding. A failed link is noticed after max age 20 s, so a topology change costs 30-50 s. During listening and learning no user traffic passes.
- RSTP (802.1w) states are discarding, learning, forwarding. Every switch originates its own BPDUs every 2 s; three missed hellos (6 s) means the neighbor is gone. On point-to-point full-duplex links a proposal/agreement handshake lets a new designated port go forwarding in well under a second, and an alternate port takes over immediately when the root port fails. Convergence is sub-second to a few seconds. On a topology change the switch flushes its MAC table so traffic reflows. MST (802.1s) maps many VLANs to a few instances; Cisco PVST+/Rapid-PVST+ runs one instance per VLAN.
- PortFast / edge port: a port facing a host skips listening and learning so the host gets link, DHCP and 802.1X in time. The risk is someone plugging a switch into it. BPDU guard error-disables an edge port the moment a BPDU arrives, which is the correct intent: the loop never forms. Root guard on downstream ports refuses a superior BPDU so a rogue switch cannot become root. Loop guard keeps a port blocked if BPDUs stop arriving on a unidirectional link.
- Why fabrics avoid STP. STP gives you one active path and blocks the rest, so half your uplinks sit idle, convergence is tens of seconds (or seconds with RSTP), failure domains are huge and a single bad BPDU reshapes the whole tree. A leaf-spine fabric instead makes every switch-to-switch link a routed point-to-point link on a /31 (or IPv6 link-local, unnumbered), runs BGP or OSPF, and uses ECMP across all uplinks at once. IP has a TTL so a routing loop is self-limiting, and convergence is a route withdrawal, not a tree rebuild. Layer 2 stops at the rack; if an application needs L2 across racks it gets a VXLAN overlay with EVPN, never a stretched VLAN.
SW2# show spanning-tree vlan 10
VLAN0010
Spanning tree enabled protocol rstp
Root ID Priority 4106 (priority 4096 sys-id-ext 10)
Address 0011.2233.4455
Cost 4
Port 48 (GigabitEthernet1/0/48)
Bridge ID Priority 32778 (priority 32768 sys-id-ext 10)
Address 00aa.bbcc.dd01
Interface Role Sts Cost Prio.Nbr Type
Gi1/0/47 Altn BLK 4 128.47 P2p <-- the blocked redundant uplink
Gi1/0/48 Root FWD 4 128.48 P2p
Gi1/0/1 Desg FWD 4 128.1 P2p Edge
LAG and LACP
A link aggregation group (EtherChannel, port-channel, bond) bundles 2-8 physical links of the same speed and duplex into one logical interface. STP sees one link, so nothing is blocked, and the bundle survives a member failure in milliseconds. LACP (802.3ad, now 802.1AX) is the negotiation protocol: each side sends LACPDUs every 1 s (fast) or 30 s (slow) with its system ID and a key; links that agree on partner and key join the bundle. Modes: active sends LACPDUs, passive only answers; at least one side must be active. A static LAG (mode on) has no protocol and no way to detect that one member is cabled to the wrong switch, so it can silently blackhole a fraction of flows. Use LACP.
- Hashing per flow. The switch computes a hash over header fields (source and destination MAC, IP, L4 ports; the 5-tuple is the usual default) and picks a member link. Every frame of one flow hashes the same way, so a flow is never reordered. The consequence: one flow can never exceed one member's bandwidth. A 4 × 10G LAG is 40G for many flows and 10G for a single backup job. Few, fat flows polarise onto one or two members; look at per-member counters before believing "the LAG is fine, it is only 30% used".
- Requirements. Same speed, duplex, VLAN configuration and MTU on every member; mismatches are the common reason a port "will not bundle". Hosts bond too: Linux
bond0mode 802.3ad withxmit_hash_policy layer3+4. - MLAG (vPC, MC-LAG, VLT) lets two physical switches present one LACP system ID so a server or downstream switch can bundle one link to each and have both active with no STP blocking. The pair syncs MAC and ARP state over a peer link and needs a keepalive to avoid both becoming primary when the peer link dies (split brain). It is vendor-specific and fragile at scale, which is why fabrics prefer plain L3 ECMP, or EVPN multihoming when a dual-homed L2 server is unavoidable.
SW1# show etherchannel summary
Group Port-channel Protocol Ports
1 Po1(SU) LACP Gi1/0/47(P) Gi1/0/48(P) S = L2, U = in use, P = bundled
Gi1/0/46(I) I = standalone: LACP never agreed, check the far end
$ cat /proc/net/bonding/bond0 | grep -E 'Mode|Hash|Slave Interface|MII Status'
Bonding Mode: IEEE 802.3ad Dynamic link aggregation
Transmit Hash Policy: layer3+4 (1)
Slave Interface: eth0 MII Status: up
Slave Interface: eth1 MII Status: up
Design sketch: enterprise LAN for 100 hosts that scales to 1000
A five-minute whiteboard answer
- 100 hosts. Two or three 48-port access switches and a collapsed core: a pair of layer-3 switches that are distribution and core at once. Each access switch has one uplink to each core switch, bundled where possible. VLANs by function, each a /24 with its SVI on the core pair: users, servers, printers, voice, Wi-Fi, management, guest. First-hop redundancy (HSRP/VRRP) for every SVI, DHCP relay on the SVIs to a central server, default route to the firewall.
- Layer 2 discipline. Rapid-PVST+ or MST with the core pair as root and backup root at fixed priorities, PortFast plus BPDU guard on every host port, root guard toward access, native VLAN set to an unused VLAN, allowed-VLAN lists pruned, no VLAN 1 for users. Alternatively run MLAG on the core pair so access uplinks are one LAG and nothing blocks.
- Scaling to 1000. Add access switches per closet and move to a three-tier shape or, better, routed access: each access switch is an L3 device, VLANs live only inside one switch, every uplink is a routed /31, OSPF (one area, or an area per building) between access, distribution and core, ECMP over both uplinks, no STP across uplinks at all. Subnet per closet, summarised toward the core. Keep a single broadcast domain under about 250 hosts.
- Security. 802.1X with dynamic VLAN assignment, DHCP snooping, dynamic ARP inspection, port security on static ports, ACLs between VLANs at the SVI or a firewall for servers, separate management VLAN and out-of-band access to every switch.
- Management. Configuration from templates in version control pushed by automation, not by hand; SNMP or streaming telemetry to a monitoring system, syslog, NTP, NetFlow/sFlow for traffic visibility, LLDP so the topology can be discovered.
- Bandwidth. 1G or mGig to hosts, 10G or 25G uplinks, roughly 20:1 oversubscription at access to distribution and 4:1 toward the core, QoS classes for voice and video marked at the edge (PCP and DSCP). State the numbers; interviewers want to hear you size the uplinks.
Interview questions
1. How does a switch build its MAC table, and what happens to a frame for an unknown destination?
It records the source MAC and ingress port of every frame it receives, and ages entries out after about 300 s of silence. A frame to an unknown unicast MAC is flooded out every other port in that VLAN; the reply teaches the switch where that MAC lives. Broadcasts are always flooded. Follow-up: a MAC learned on two ports alternately means a loop or a duplicate MAC.2. What is the difference between a collision domain and a broadcast domain?
A collision domain is the set of devices sharing one medium; on a full-duplex switched port it is just the two ends, so collisions are a hub-era problem. A broadcast domain is everything a broadcast frame reaches: one VLAN, bounded by a router. Switches split collision domains; routers and VLANs split broadcast domains.3. What is a VLAN and why would you use one?
A VLAN is a separate broadcast domain inside one physical switch, with its own MAC table and usually its own IP subnet. You use VLANs to limit broadcast scope, group hosts by function or security policy regardless of physical port, and keep management or voice traffic apart. Traffic between VLANs must go through a router or an L3 switch.4. Describe the 802.1Q tag.
Four bytes inserted after the source MAC: a 16-bit TPID of 0x8100 that says a tag follows, then 3 bits of priority, 1 drop-eligible bit, and a 12-bit VLAN ID. 12 bits gives 4096 values; 0 and 4095 are reserved, so 4094 usable VLANs. QinQ stacks a second tag with TPID 0x88a8.5. Access port vs trunk port, and what is the native VLAN?
An access port carries one VLAN and the frames are untagged on the wire; the switch assigns the VLAN. A trunk carries many VLANs with each frame tagged. The native VLAN is the one VLAN a trunk sends untagged, default VLAN 1. If the two ends disagree, untagged frames land in different VLANs on each side and traffic leaks, so you either match them or tag the native VLAN.6. How do two hosts in different VLANs communicate?
Through a layer-3 device. Either a router-on-a-stick, one trunk to a router with a subinterface per VLAN, where every inter-VLAN packet crosses that one link twice, or an SVI per VLAN on a layer-3 switch, routed in hardware at line rate. The host sees only its gateway; it ARPs for the SVI address and the switch routes and rewrites the MAC like any router.7. Why are layer-2 loops so dangerous?
Ethernet frames have no TTL, so a broadcast on a looped topology is flooded forever and multiplies at every switch. Within seconds you have a broadcast storm saturating links, MAC tables flapping because the same source appears on two ports, and switch CPUs pinned. Unlike a routing loop, nothing times the frames out. STP blocks redundant ports to prevent it.8. How is the root bridge elected and how does a switch pick its root port?
Every switch starts claiming to be root in its BPDUs; the lowest bridge ID wins, and bridge ID is priority then MAC, so with default priorities the oldest switch wins. You set a low priority on the switch you want. Each other switch then chooses one root port: the lowest total path cost to the root, with ties broken by the neighbor's bridge ID and then the neighbor's port ID. On each segment the switch with the lowest root cost has the designated port; everything else blocks.9. Why does RSTP converge faster than classic STP?
Classic STP waits for max age (20 s) to notice a failure and then spends 15 s listening and 15 s learning. RSTP has every switch send its own hellos so loss is detected in 6 s, keeps alternate ports precomputed so a root port failure is an instant switchover, and uses a proposal/agreement handshake on point-to-point links instead of timers. Result: sub-second to a few seconds instead of 30-50.10. What do PortFast and BPDU guard do, and why both?
PortFast makes a host-facing port forward immediately instead of waiting 30 s, so DHCP and boot work. That is unsafe if someone plugs in a switch. BPDU guard shuts the port the moment a BPDU arrives, so you get the speed without the loop risk. Always pair them on edge ports.11. Why do data-center fabrics not run STP?
STP wastes every redundant link by blocking it, converges slowly, and makes one layer-2 domain a single failure domain. A leaf-spine fabric runs IP on every link, a /31 per link, with BGP or OSPF and ECMP so all uplinks carry traffic, and IP TTL makes loops self-limiting. Layer 2 ends at the rack; if a workload needs L2 adjacency across racks you build a VXLAN/EVPN overlay rather than stretching a VLAN.12. You bundle four 10G links with LACP. Why does one file transfer still get only 10G?
The LAG hashes each flow by header fields and pins it to one member so frames stay in order. A single TCP flow therefore uses one 10G link; the 40G is only reachable with many flows. Fixes are more parallel flows, a faster member link, or hashing on fields that actually differ between your flows. Follow-up: what does LACP give you over a static bundle? Detection of a miswired or dead member so it is removed instead of blackholing its share of flows.Traps
- Saying a switch learns from destination MACs, or that it "knows" IPs. It learns from sources only and never reads the IP header.
- Saying VLANs are "a security feature" without naming the mechanism: separate broadcast domains forced through a router or firewall. Native VLAN mismatches and double tagging are how VLANs leak.
- Claiming 4096 VLANs are usable. 0 and 4095 are reserved; 4094 are usable, and many platforms reserve 1002-1005 too.
- Saying STP "load balances" or "uses both links". It blocks all but one path per VLAN. Per-VLAN STP across two uplinks is a manual approximation, LAG or routed ECMP is the real answer.
- Forgetting the extended system ID when quoting a priority: a Cisco switch shows 32778 for VLAN 10, which is 32768 plus 10.
- Saying a LAG "doubles the bandwidth" for a flow. It multiplies aggregate bandwidth across flows; one flow is still capped at one member.
- Confusing the layer-2 loop problem with a routing loop. Routing loops die at TTL 0; frame loops never die.
- Describing router-on-a-stick as the normal way to route between VLANs in a modern network. It is the teaching example; SVIs on an L3 switch are what you would deploy.
Scenario
A new rack gets a third VLAN, 30, for storage. You create VLAN 30 on both the access switch SW1 and the core pair and configure switchport access vlan 30 on the storage ports. The storage hosts on SW1 can ping each other but get 169.254 addresses from DHCP and cannot reach the gateway 10.1.30.1. VLAN 10 and 20 hosts on the same switch are fine. Walk through it.
Expected reasoning
- Scope it. Storage hosts reach each other, so link, cabling, NICs, and VLAN 30 on SW1's local ports all work. They cannot reach anything beyond SW1 in that VLAN, while VLANs 10 and 20 cross the uplink fine. The break is specifically VLAN 30 between SW1 and the core: the trunk or the SVI.
- Check the trunk from SW1.
show interfaces trunk: is 30 in "VLANs allowed on trunk" and in "VLANs allowed and active in management domain"? The common bug isswitchport trunk allowed vlan 10,20set years ago; adding VLAN 30 to the database does not add it to the trunk. Fix withswitchport trunk allowed vlan add 30(the word add; without it you replace the list and take down VLANs 10 and 20). - Check the other end. Same command on the core. The VLAN must exist and be allowed on its side of the trunk too. Check the native VLAN matches on both ends while you are there. If the trunk is a port-channel, check the channel, not just a member.
- Check STP.
show spanning-tree vlan 30on both ends: is the uplink forwarding for this VLAN? With per-VLAN STP a new VLAN takes up to 30 s to go forwarding on a non-edge port; a blocked state here with 10 and 20 forwarding points at a per-VLAN cost or priority tweak. - Check the gateway. On the core:
show ip interface brief | include Vlan30. The SVI must be up/up with 10.1.30.1, and an SVI is only up if the VLAN exists and at least one port in it is up. Confirmip helper-addressis on the Vlan30 SVI, otherwise the hosts will ARP the gateway fine but still get 169.254. - Confirm with evidence. After the fix, the core MAC table shows storage MACs in VLAN 30 on the trunk, the hosts
ip neighshows 10.1.30.1 REACHABLE, and adhclient -vrun shows DORA completing. Say the root cause in one line: "VLAN 30 was created but not added to the allowed list on the SW1 uplink, so its frames were dropped at the trunk."
← 9 · Strings, two pointers and sliding windows · all topics · 11 · Stacks, queues, intervals and 2D grids →