← all topics

10 · Switching: MAC learning, VLANs, STP and LAG Network

Flood/learn/forward · 802.1Q trunks · root bridge and blocked ports · LACP and hashing

Why it matters for NPE. Layer 2 questions (MAC tables, VLAN mismatches, loops) appear in Menlo Park reports and are where many software-leaning candidates are weakest.

Primer: a switch does three things to a frame

A switch learns the source MAC of every frame and the port it arrived on, forwards a frame whose destination MAC it knows out that one port, and floods everything else (unknown unicast, broadcast, most multicast) out every other port in the same VLAN. That is the whole data plane. VLANs split one switch into several of these broadcast domains. STP exists because flooding plus a physical loop is fatal. LAG exists because you want two links between switches without STP blocking one of them.

The sentence to open with: "Learn on source, forward on destination, flood when unknown. Everything else in switching is a patch for what flooding breaks."

Watch

Everything Switches do - Part 1 - Networking Fundamentals - Lesson 4Practical Networking · 11:37

Learn, flood, forward and filter, shown frame by frame as the MAC table fills. Eleven minutes; watch the whole thing and redraw the table on paper afterwards.

Free CCNA | VLANs (Part 1) | Day 16 | CCNA 200-301 Complete CourseJeremy's IT Lab · 23:44

Watch 0:00-13:05: 2:02 broadcast domains, 4:38 what a VLAN is, 6:22 segmenting at layer 3, 10:07 segmenting at layer 2. The rest is Cisco CLI.

Free CCNA | VLANs (Part 2) | Day 17 | CCNA 200-301 Complete CourseJeremy's IT Lab · 40:01

Watch 4:26-16:53: 4:26 trunk ports, 8:15 the 802.1Q tag, 9:12 TPID, 10:14 PCP, 11:02 VID, 12:22 VLAN ranges, 13:44 native VLAN. 27:00 router-on-a-stick if inter-VLAN routing is new.

Free CCNA | Spanning Tree Protocol (Part 1) | Day 20 | CCNA 200-301 Complete CourseJeremy's IT Lab · 38:39

Watch 1:47-33:17: 4:15 loops and broadcast storms, 11:53 BPDUs and root election, 21:00 root port by cost, 24:21 by neighbor bridge ID, 26:48 by port ID, 29:13 blocking, 31:20 process summary.

Free CCNA | EtherChannel | Day 23 | CCNA 200-301 Complete CourseJeremy's IT Lab · 41:33

Watch 1:48-17:57: 1:48 why bundle links, 8:59 load balancing by per-flow hash (why one flow never exceeds one member), 15:30 PAgP vs LACP vs static. Skip config from 17:57.

Free CCNA | Rapid Spanning Tree Protocol | Day 22 | CCNA 200-301 Complete CourseJeremy's IT Lab · 43:01

Only 2:19-7:21, the STP vs RSTP comparison table. The port role and state tables below cover the rest.

VLANs, Tagging, Trunking, VxLAN, & Native VLANPowerCert Animated Videos · 9:30

Optional refresher. 2:49 tagging and trunking, 4:40 native VLAN, 6:31 VXLAN as the bridge to data-center overlay vocabulary.

Learn, forward, flood: a worked MAC table

A (aa) B (bb) C (cc) │p1 │p2 │p3 ┌─┴─────────────┴─────────────┴─┐ │ Switch │ MAC table starts EMPTY └────────────────────────────────┘ 1. A sends to B (dst bb). Switch learns aa→p1. bb unknown → FLOOD out p2, p3. table: aa p1 2. B replies to A (dst aa). Switch learns bb→p2. aa known → FORWARD out p1 only. C sees nothing. table: aa p1 | bb p2 3. C broadcasts an ARP request (dst ff:ff:ff:ff:ff:ff). Learns cc→p3. Broadcast → FLOOD out p1, p2. table: aa p1 | bb p2 | cc p3 4. A sends to B again. Both known → FORWARD p1→p2. If dst were on the same port it came in: FILTER (drop). 5. Nothing from C for 300 s → cc→p3 AGES OUT. Next frame to cc is flooded until C speaks again.
Switch# show mac address-table
Vlan  Mac Address      Type     Ports
  10  aaaa.aaaa.aaaa   DYNAMIC  Gi1/0/1
  10  bbbb.bbbb.bbbb   DYNAMIC  Gi1/0/2
  20  cccc.cccc.cccc   DYNAMIC  Gi1/0/3
  10  0000.0c9f.f00a   DYNAMIC  Gi1/0/48     <-- gateway MAC (HSRP) learned on the uplink

$ bridge fdb show br0 | grep -v permanent          # Linux bridge, same idea
aa:aa:aa:aa:aa:aa dev eth1 vlan 10 master br0

VLANs and 802.1Q

Untagged frame: | dst MAC | src MAC | EtherType | payload | FCS | 802.1Q tagged: | dst MAC | src MAC | TPID 0x8100 | PCP(3) DEI(1) VID(12) | EtherType | payload | FCS | <------------ 4-byte tag -----------> TPID 16 bits 0x8100 says "a tag follows" (0x88a8 for the outer tag in QinQ / 802.1ad) PCP 3 bits priority 0-7 (CoS), used for QoS at layer 2 DEI 1 bit drop eligible (formerly CFI) VID 12 bits VLAN 0-4095. 0 = priority tag only, 4095 reserved, so 1-4094 usable = 4094 VLANs Max tagged frame 1522 bytes (1518 + 4)
Access portTrunk port
Carriesone VLAN, frames untagged on the wiremany VLANs, each frame tagged with its VID
Connects toa host, server, printer, phone (voice VLAN is the one exception, tagged alongside the untagged data VLAN)another switch, a router doing inter-VLAN routing, a hypervisor
Ingressswitch assigns the port's VLANswitch reads the VID; untagged frames go to the native VLAN
Egresstag removedtag kept, except frames in the native VLAN which leave untagged
Configswitchport mode access, switchport access vlan 10switchport mode trunk, switchport trunk allowed vlan 10,20,30, switchport trunk native vlan 999
$ tcpdump -eni eth0 vlan
aa:aa:aa:aa:aa:aa > ff:ff:ff:ff:ff:ff, ethertype 802.1Q (0x8100), length 46: vlan 10, p 0, ethertype ARP, Request who-has 10.1.10.1 tell 10.1.10.50

SW1# show interfaces trunk
Port      Mode   Encapsulation  Status    Native vlan
Gi1/0/48  on     802.1q         trunking  999
Port      Vlans allowed on trunk
Gi1/0/48  10,20          <-- VLAN 30 missing here is the whole bug

Spanning Tree: loops are fatal at layer 2

An Ethernet frame has no TTL. Connect two switches with two cables and send one broadcast: each switch floods it out the other link, the other switch floods it back, forever, at line rate. Within seconds the links are saturated (a broadcast storm), every switch sees the same source MAC arriving on two ports and rewrites its table on each frame (MAC flapping), unicast traffic is misdirected, and CPUs melt processing broadcasts. Hosts receive thousands of copies of every ARP. The only exit is pulling a cable. STP's job is to make the physical loop logically a tree by blocking ports.

BPDU (sent every hello = 2 s to 01:80:c2:00:00:00) root bridge ID | root path cost | sender bridge ID | sender port ID | hello 2 | max age 20 | fwd delay 15 bridge ID = priority (4 bits, default 32768, steps of 4096) + ext. system ID (12 bits, the VLAN) + MAC LOWEST bridge ID wins the root. Equal priority everywhere means the OLDEST switch (lowest MAC) is root. Set priority 4096 on the switch you want as root (and 8192 on the backup); never leave it to chance.
Port roleHow chosenState
Root portOne per non-root switch: the port with the lowest cost to the root. Ties: lowest neighbor bridge ID, then lowest neighbor port ID.forwarding
Designated portOne per segment: the port on the switch with the lowest root cost on that segment (same tie-breaks). All root bridge ports are designated.forwarding
Non-designated (802.1D) / Alternate, Backup (RSTP)Everything else. Alternate = a second path to the root, Backup = a second port onto a segment I already serve.blocking (RSTP: discarding). Still receives BPDUs so it can wake up.
Link speed802.1D short costRSTP/MST long cost
10 Mbit/s1002 000 000
100 Mbit/s19200 000
1 Gbit/s420 000
10 Gbit/s22 000
100 Gbit/s1 (useless)200
SW2# show spanning-tree vlan 10
VLAN0010
  Spanning tree enabled protocol rstp
  Root ID    Priority    4106      (priority 4096 sys-id-ext 10)
             Address     0011.2233.4455
             Cost        4
             Port        48 (GigabitEthernet1/0/48)
  Bridge ID  Priority    32778     (priority 32768 sys-id-ext 10)
             Address     00aa.bbcc.dd01
Interface        Role Sts Cost      Prio.Nbr Type
Gi1/0/47         Altn BLK 4         128.47   P2p     <-- the blocked redundant uplink
Gi1/0/48         Root FWD 4         128.48   P2p
Gi1/0/1          Desg FWD 4         128.1    P2p Edge

LAG and LACP

A link aggregation group (EtherChannel, port-channel, bond) bundles 2-8 physical links of the same speed and duplex into one logical interface. STP sees one link, so nothing is blocked, and the bundle survives a member failure in milliseconds. LACP (802.3ad, now 802.1AX) is the negotiation protocol: each side sends LACPDUs every 1 s (fast) or 30 s (slow) with its system ID and a key; links that agree on partner and key join the bundle. Modes: active sends LACPDUs, passive only answers; at least one side must be active. A static LAG (mode on) has no protocol and no way to detect that one member is cabled to the wrong switch, so it can silently blackhole a fraction of flows. Use LACP.

SW1# show etherchannel summary
Group  Port-channel  Protocol    Ports
1      Po1(SU)       LACP        Gi1/0/47(P)  Gi1/0/48(P)      S = L2, U = in use, P = bundled
                                 Gi1/0/46(I)                    I = standalone: LACP never agreed, check the far end

$ cat /proc/net/bonding/bond0 | grep -E 'Mode|Hash|Slave Interface|MII Status'
Bonding Mode: IEEE 802.3ad Dynamic link aggregation
Transmit Hash Policy: layer3+4 (1)
Slave Interface: eth0    MII Status: up
Slave Interface: eth1    MII Status: up

Design sketch: enterprise LAN for 100 hosts that scales to 1000

A five-minute whiteboard answer
  1. 100 hosts. Two or three 48-port access switches and a collapsed core: a pair of layer-3 switches that are distribution and core at once. Each access switch has one uplink to each core switch, bundled where possible. VLANs by function, each a /24 with its SVI on the core pair: users, servers, printers, voice, Wi-Fi, management, guest. First-hop redundancy (HSRP/VRRP) for every SVI, DHCP relay on the SVIs to a central server, default route to the firewall.
  2. Layer 2 discipline. Rapid-PVST+ or MST with the core pair as root and backup root at fixed priorities, PortFast plus BPDU guard on every host port, root guard toward access, native VLAN set to an unused VLAN, allowed-VLAN lists pruned, no VLAN 1 for users. Alternatively run MLAG on the core pair so access uplinks are one LAG and nothing blocks.
  3. Scaling to 1000. Add access switches per closet and move to a three-tier shape or, better, routed access: each access switch is an L3 device, VLANs live only inside one switch, every uplink is a routed /31, OSPF (one area, or an area per building) between access, distribution and core, ECMP over both uplinks, no STP across uplinks at all. Subnet per closet, summarised toward the core. Keep a single broadcast domain under about 250 hosts.
  4. Security. 802.1X with dynamic VLAN assignment, DHCP snooping, dynamic ARP inspection, port security on static ports, ACLs between VLANs at the SVI or a firewall for servers, separate management VLAN and out-of-band access to every switch.
  5. Management. Configuration from templates in version control pushed by automation, not by hand; SNMP or streaming telemetry to a monitoring system, syslog, NTP, NetFlow/sFlow for traffic visibility, LLDP so the topology can be discovered.
  6. Bandwidth. 1G or mGig to hosts, 10G or 25G uplinks, roughly 20:1 oversubscription at access to distribution and 4:1 toward the core, QoS classes for voice and video marked at the edge (PCP and DSCP). State the numbers; interviewers want to hear you size the uplinks.

Interview questions

1. How does a switch build its MAC table, and what happens to a frame for an unknown destination?It records the source MAC and ingress port of every frame it receives, and ages entries out after about 300 s of silence. A frame to an unknown unicast MAC is flooded out every other port in that VLAN; the reply teaches the switch where that MAC lives. Broadcasts are always flooded. Follow-up: a MAC learned on two ports alternately means a loop or a duplicate MAC.
2. What is the difference between a collision domain and a broadcast domain?A collision domain is the set of devices sharing one medium; on a full-duplex switched port it is just the two ends, so collisions are a hub-era problem. A broadcast domain is everything a broadcast frame reaches: one VLAN, bounded by a router. Switches split collision domains; routers and VLANs split broadcast domains.
3. What is a VLAN and why would you use one?A VLAN is a separate broadcast domain inside one physical switch, with its own MAC table and usually its own IP subnet. You use VLANs to limit broadcast scope, group hosts by function or security policy regardless of physical port, and keep management or voice traffic apart. Traffic between VLANs must go through a router or an L3 switch.
4. Describe the 802.1Q tag.Four bytes inserted after the source MAC: a 16-bit TPID of 0x8100 that says a tag follows, then 3 bits of priority, 1 drop-eligible bit, and a 12-bit VLAN ID. 12 bits gives 4096 values; 0 and 4095 are reserved, so 4094 usable VLANs. QinQ stacks a second tag with TPID 0x88a8.
5. Access port vs trunk port, and what is the native VLAN?An access port carries one VLAN and the frames are untagged on the wire; the switch assigns the VLAN. A trunk carries many VLANs with each frame tagged. The native VLAN is the one VLAN a trunk sends untagged, default VLAN 1. If the two ends disagree, untagged frames land in different VLANs on each side and traffic leaks, so you either match them or tag the native VLAN.
6. How do two hosts in different VLANs communicate?Through a layer-3 device. Either a router-on-a-stick, one trunk to a router with a subinterface per VLAN, where every inter-VLAN packet crosses that one link twice, or an SVI per VLAN on a layer-3 switch, routed in hardware at line rate. The host sees only its gateway; it ARPs for the SVI address and the switch routes and rewrites the MAC like any router.
7. Why are layer-2 loops so dangerous?Ethernet frames have no TTL, so a broadcast on a looped topology is flooded forever and multiplies at every switch. Within seconds you have a broadcast storm saturating links, MAC tables flapping because the same source appears on two ports, and switch CPUs pinned. Unlike a routing loop, nothing times the frames out. STP blocks redundant ports to prevent it.
8. How is the root bridge elected and how does a switch pick its root port?Every switch starts claiming to be root in its BPDUs; the lowest bridge ID wins, and bridge ID is priority then MAC, so with default priorities the oldest switch wins. You set a low priority on the switch you want. Each other switch then chooses one root port: the lowest total path cost to the root, with ties broken by the neighbor's bridge ID and then the neighbor's port ID. On each segment the switch with the lowest root cost has the designated port; everything else blocks.
9. Why does RSTP converge faster than classic STP?Classic STP waits for max age (20 s) to notice a failure and then spends 15 s listening and 15 s learning. RSTP has every switch send its own hellos so loss is detected in 6 s, keeps alternate ports precomputed so a root port failure is an instant switchover, and uses a proposal/agreement handshake on point-to-point links instead of timers. Result: sub-second to a few seconds instead of 30-50.
10. What do PortFast and BPDU guard do, and why both?PortFast makes a host-facing port forward immediately instead of waiting 30 s, so DHCP and boot work. That is unsafe if someone plugs in a switch. BPDU guard shuts the port the moment a BPDU arrives, so you get the speed without the loop risk. Always pair them on edge ports.
11. Why do data-center fabrics not run STP?STP wastes every redundant link by blocking it, converges slowly, and makes one layer-2 domain a single failure domain. A leaf-spine fabric runs IP on every link, a /31 per link, with BGP or OSPF and ECMP so all uplinks carry traffic, and IP TTL makes loops self-limiting. Layer 2 ends at the rack; if a workload needs L2 adjacency across racks you build a VXLAN/EVPN overlay rather than stretching a VLAN.
12. You bundle four 10G links with LACP. Why does one file transfer still get only 10G?The LAG hashes each flow by header fields and pins it to one member so frames stay in order. A single TCP flow therefore uses one 10G link; the 40G is only reachable with many flows. Fixes are more parallel flows, a faster member link, or hashing on fields that actually differ between your flows. Follow-up: what does LACP give you over a static bundle? Detection of a miswired or dead member so it is removed instead of blackholing its share of flows.

Traps

Scenario

A new rack gets a third VLAN, 30, for storage. You create VLAN 30 on both the access switch SW1 and the core pair and configure switchport access vlan 30 on the storage ports. The storage hosts on SW1 can ping each other but get 169.254 addresses from DHCP and cannot reach the gateway 10.1.30.1. VLAN 10 and 20 hosts on the same switch are fine. Walk through it.

Expected reasoning
  1. Scope it. Storage hosts reach each other, so link, cabling, NICs, and VLAN 30 on SW1's local ports all work. They cannot reach anything beyond SW1 in that VLAN, while VLANs 10 and 20 cross the uplink fine. The break is specifically VLAN 30 between SW1 and the core: the trunk or the SVI.
  2. Check the trunk from SW1. show interfaces trunk: is 30 in "VLANs allowed on trunk" and in "VLANs allowed and active in management domain"? The common bug is switchport trunk allowed vlan 10,20 set years ago; adding VLAN 30 to the database does not add it to the trunk. Fix with switchport trunk allowed vlan add 30 (the word add; without it you replace the list and take down VLANs 10 and 20).
  3. Check the other end. Same command on the core. The VLAN must exist and be allowed on its side of the trunk too. Check the native VLAN matches on both ends while you are there. If the trunk is a port-channel, check the channel, not just a member.
  4. Check STP. show spanning-tree vlan 30 on both ends: is the uplink forwarding for this VLAN? With per-VLAN STP a new VLAN takes up to 30 s to go forwarding on a non-edge port; a blocked state here with 10 and 20 forwarding points at a per-VLAN cost or priority tweak.
  5. Check the gateway. On the core: show ip interface brief | include Vlan30. The SVI must be up/up with 10.1.30.1, and an SVI is only up if the VLAN exists and at least one port in it is up. Confirm ip helper-address is on the Vlan30 SVI, otherwise the hosts will ARP the gateway fine but still get 169.254.
  6. Confirm with evidence. After the fix, the core MAC table shows storage MACs in VLAN 30 on the trunk, the hosts ip neigh shows 10.1.30.1 REACHABLE, and a dhclient -v run shows DORA completing. Say the root cause in one line: "VLAN 30 was created but not added to the allowed list on the SW1 uplink, so its frames were dropped at the trunk."

← 9 · Strings, two pointers and sliding windows · all topics · 11 · Stacks, queues, intervals and 2D grids →