21 · Network interview script, behavioral and 'Why Meta' Network
Two protocols in depth with pros/cons and comparison (BGP and TCP, OSPF as the foil) · automation stories · behavioral questions · design-round preview · questions to ask
Why it matters for NPE. Meta's own guide says to prepare at least two protocols in depth with their trade-offs. The loop adds behavioral and, for some roles, a network-design round. Prepare three stories from your own work and a real answer to 'Why Meta'.
Primer: the network interview is a conversation you can steer
The network round is 45 minutes of talking, no code. Meta's own screen guide asks for at least two protocols in depth, with pros and cons and a comparison to other protocols. Candidate reports agree on what gets probed: BGP most often, TCP second, the life of a packet, OSPF mostly as "why not OSPF", and one troubleshooting scenario. The interviewer goes deep on whatever you claim to know. The most common failure story is "I prepared breadth, they asked BGP depth". So decide now: BGP is your deepest protocol, TCP is your second, OSPF is the comparison foil. Everything on this page serves that plan.
Watch
A rapid-fire mock whose questions map almost one to one onto Meta's "explain it in depth" format. Use it as a drill: pause at each chapter title, answer aloud, then compare. Chapters: 1:31 TCP header, 2:27 flags, 4:34 handshake, 5:11 window scaling, 5:40 two PCs on a switch and the flow of an ICMP packet, 6:34 how ARP is sent, 11:50 DNS, 13:00 OSPF summarisation, 15:37 LSA types, 17:52 DR/BDR, 21:42 root bridge placement. Ignore the Cisco-only bits (port security modes).
A former Meta engineering manager and hiring-committee member on what the behavioral round scores. Watch 12:43 understanding interviewer signals and 18:11 presenting your experience, then 25:35 common mistakes and red flags, 37:00 team match and asking questions.
The minimum to talk credibly about "how would you automate this across 500 switches": 2:31 ConnectHandler, 3:55 send_command, 5:25 closing the session. Pair with the automation section below and topic 15.
Meta's own talk. Watch it again before the interview; it is the backbone of your "Why Meta" answer and of "why BGP instead of OSPF in the data center".
The PE loop stages in order (0:40 screen, 1:56 online assessment, 2:34 coding, 3:49 systems, 5:42 design). Generic PE, not NPE, but accurate on structure. Watch at 1.5x.
Reading beats video for this topic: the official guide's "Final tips" page (be yourself, research the role, it is ok not to know, prepare questions), Meta's 2021 post on BGP in the data center, and the F16/Minipack post. Both posts are linked from topic 14 and topic 16.
The 45 minutes as a timeline
| Minutes | What happens | How you steer |
|---|---|---|
| 0-5 | Introductions. "Tell me about yourself and your networking background." | Ninety seconds, not five minutes. Degree and focus, the internship in one sentence (what the team did, what you built), then: "The protocols I know best are BGP and TCP; I have also worked with OSPF, so I can compare them." You have just set the agenda. |
| 5-20 | Protocol depth. The interviewer picks one protocol and follows up for 10-15 minutes. | Run the deep-dive script below. Pause after each block. If they ask "anything else?", offer the next block rather than trivia. |
| 20-30 | Second protocol or the comparison. "How does that compare with OSPF?" or "Walk me through TCP." | Give the 45-second OSPF comparison or the TCP script. Name the trade-off, not just the difference. |
| 30-40 | Scenario. Life of a packet, unreachable server, latency and loss, or a BGP peer that will not come up. | Say the layers you will check and in what order before you check anything. See topic 19. |
| 40-45 | Your questions. | Two prepared questions, one about the team's day to day, one technical. Not "what is the culture like". |
If the interviewer opens with "pick your favourite L2 and L3 protocol", answer: "At L3, BGP, because it is where policy and scale meet. At L2, I would pick Ethernet switching with STP, because almost every outage story I have seen starts there." Then wait.
The protocol deep-dive script
Use the same five blocks for any protocol. Each block is a natural stopping point where the interviewer can redirect.
| Block | Length | Content |
|---|---|---|
| 1. Overview | 60 s | What problem it solves, where it sits in the stack, what it runs over, one sentence of mental model. |
| 2. How it works | 3-5 min | Messages, states, timers, tables. Name the header fields or attributes. Walk one exchange end to end. |
| 3. What breaks and how you'd see it | 2-3 min | Three failure modes, each with the command or capture pattern that reveals it. |
| 4. Trade-offs and comparison | 1-2 min | What it is bad at. The protocol you would pick instead and when. |
| 5. Where you've used it | 30 s | A real place you configured, simulated, debugged or read about it. True details only. |
BGP, filled in (full detail in topic 14)
- Overview. "BGP is a path-vector routing protocol between autonomous systems. It exchanges prefixes with attributes over a TCP session on port 179, prevents loops with the AS_PATH, and picks one best path per prefix with a strict attribute comparison that operators steer with policy."
- How it works. Peers are configured by IP and AS, never discovered. FSM: Idle, Connect, Active, OpenSent, OpenConfirm, Established. Messages: OPEN (AS, hold time, router ID, capabilities), KEEPALIVE (every 60 s, hold 180 s), UPDATE (prefixes plus attributes, or withdrawals), NOTIFICATION (error, then close). Tables: RIB-in per peer, Loc-RIB, RIB-out per peer. eBGP prepends the AS and rewrites NEXT_HOP; iBGP changes neither and will not re-advertise iBGP-learned routes, hence full mesh or route reflectors. Best path: next hop reachable, then weight, local preference, locally originated, shortest AS_PATH, origin, MED, eBGP over iBGP, IGP metric to next hop, then tie-breakers.
- What breaks. Stuck in Active: TCP cannot complete, so ping, port 179, ACLs, multihop, update source. Flapping: hold timer expiry from loss or CPU, visible as the last NOTIFICATION in
show bgp neighbor. "Route is there but not best": NEXT_HOP unreachable or a local-pref mismatch, visible inshow bgp ipv4 unicast <prefix>. Route leak: a full table where a few prefixes were expected, caught by max-prefix. - Trade-offs. Slow convergence by default (use BFD or short timers), no idea of latency or bandwidth, configuration heavy. In exchange: policy per neighbour, topology hiding, Internet scale, and a small eBGP-only design that works for a Clos fabric with a private ASN per tier.
- Where you've used it. Say exactly what you did: a lab, a simulation, a course, a troubleshooting session. If your internship involved virtual networks or simulated topologies, describe what you actually built and observed. Do not upgrade "I watched a peer come up in a lab" into "I ran production BGP".
TCP, filled in (full detail in topic 8)
- Overview. "TCP turns IP's best-effort datagrams into a reliable, ordered byte stream between two ports. It does that with sequence numbers, acknowledgements, retransmission timers, a receiver-advertised window for flow control and a sender-side congestion window for congestion control."
- How it works. Header: source and destination port, sequence number, acknowledgement number, data offset, flags (SYN, ACK, FIN, RST, PSH, URG), window, checksum, options (MSS, window scale, SACK, timestamps). Handshake: SYN, SYN-ACK, ACK; states LISTEN, SYN-SENT, SYN-RECEIVED, ESTABLISHED. Teardown: FIN from each side, FIN-WAIT, CLOSE-WAIT, LAST-ACK, TIME-WAIT. Reliability: RTO from measured RTT, fast retransmit on three duplicate ACKs, SACK to fill holes. Flow control: receive window. Congestion control: slow start, congestion avoidance, cwnd halves or resets on loss; CUBIC and BBR as modern algorithms.
- What breaks. SYN with no answer: a filter dropping packets (silence) versus no listener (RST). Retransmissions and duplicate ACKs in a capture: loss on the path. Zero window: the receiver application is not reading. Handshake succeeds but large transfers hang: MTU black hole, fix with MSS clamping or PMTUD. Many TIME_WAIT or CLOSE_WAIT sockets in
ss -tan: application not closing properly. - Trade-offs. Head-of-line blocking, handshake latency, state on both ends. UDP for DNS queries, real-time media, and RoCE-style fabrics where the application handles loss; QUIC moves reliability into user space over UDP.
- Where you've used it. Every socket you have debugged. A concrete capture you read, a connection timeout you chased, a server you tuned. Again, true details only.
OSPF as the comparison, 45 seconds
Expected follow-ups: "why can't we use OSPF for the whole data center?" (database size, flooding domain, no per-link policy, one failure domain), "why not iBGP instead of OSPF?" (iBGP still needs an IGP for next-hop reachability and needs full mesh or reflectors), "how does each avoid loops?" (SPF on a consistent database versus AS_PATH and the iBGP split-horizon rule). See topic 12.
It's okay not to know
The official guide says it directly: no one at Meta is an expert in everything, and you are encouraged to be upfront about topics you know less well. Bluffing is the thing that fails you. Not knowing is not.
| Situation | Say this | Then do this |
|---|---|---|
| You have never touched it | "I have not worked with that directly. Here is what I do know, and here is how I would reason about it." | Reason from first principles out loud: which layer, what problem it solves, what a protocol in that position must have (discovery, state, timers, failure detection). |
| You know it partially | "I know the concept but not the exact numbers. Let me check my understanding with you." | State what you are confident about, mark what you are guessing: "I believe the default is 180 seconds, but I would verify that." |
| You went blank | "Give me a second to structure this." | Take the second. Return to the five-block script. Start with the overview block even if the question was about block three. |
| The interviewer offers a hint | "Right, that points at the next hop. Let me follow that." | Take it immediately and say where it leads. Fighting a hint or ignoring it is a red flag in every interview report. |
| You realise you were wrong | "I need to correct what I said a minute ago." | Correct it in one sentence and move on. The guide's coding advice applies here too: changing your mind early is fine. |
First-principles example. Asked about IS-IS and you only know OSPF: "I know IS-IS is the other link-state IGP. So it must flood link state, build a database, run SPF and have hellos and a dead interval. The differences I remember are that it runs directly on layer 2 rather than over IP, uses levels instead of areas, and is popular in backbones because it carries IPv4 and IPv6 in one instance. I could not configure it from memory." That is a passing answer.
Automation and operations: how to talk about it
The job posting says "define and develop optimized network monitoring and automation systems", and one Meta employee described the role as "SWE plus networking". Expect at least one question that starts "how would you do this for 1000 devices". Have this vocabulary ready and use the words on purpose.
| Principle | What it means in one sentence |
|---|---|
| Source of truth | Device config is derived from a database or repo of intent (roles, links, ASNs, prefixes), never edited by hand and never read back as truth. |
| Config generation from templates | Jinja2 or similar renders per-device config from the source of truth; the template is reviewed once, the thousand outputs are not. |
| Idempotency | Running the job twice produces the same device state and the second run changes nothing. You get it by computing a diff and pushing only the diff. |
| Dry-run and diff | Show the intended change before applying it; a human or a policy checks the diff, not the script. |
| Canary, then expand | One device, verify, a small set, verify, then the fleet in waves with automatic stop on failure. |
| Rollback | Every change ships with a tested way back: commit-confirm, config snapshot, or regenerate from the previous source-of-truth version. |
| Pre- and post-checks | Capture BGP peer counts, route counts, interface errors before and after; fail the change if they differ beyond a threshold. |
| Telemetry and alerting | Stream counters and state (gNMI, SNMP, syslog) into a time series store; alert on symptoms users feel, page on things a human must fix. |
| Why scripts beat CLI at scale | Humans typing make different mistakes on each device; a script makes the same mistake everywhere, which is why canaries and diffs exist. Scripts are reviewable, repeatable, testable and leave an audit trail. |
Tools by name. netmiko: SSH to network devices, send commands, get text back. napalm: vendor-neutral getters (get_bgp_neighbors, get_interfaces) plus merge/replace config with diff and rollback. nornir: an inventory and task runner that parallelises netmiko or napalm across hosts. pyATS with genie: parsers that turn show output into dicts, plus test harnesses. TextFSM and ntc-templates for parsing when you only have text. One line on the model-driven side: NETCONF (XML over SSH) and gNMI (gRPC, streaming telemetry) move structured data described by YANG models, so you get dicts instead of screen scraping. At Meta the switches run FBOSS, a set of applications on Linux that program the ASIC and expose a Thrift API, with an in-house BGP agent deployed and tested like software. So the right sentence is: "I would pull structured data through an API rather than scrape CLI, and I would scrape CLI only where no API exists."
Five likely automation questions
1. Write a script that logs into 100 devices, runs show ip route, and prints each hostname with its best routes. Now scale it to 1000.
Say the shape first: inventory in, one worker per device, parse, aggregate, print. Sequential SSH at 2-3 seconds a device is 5 minutes for 100 and 50 minutes for 1000, so parallelise. Network I/O is the bottleneck, so threads are fine despite the GIL (concurrent.futures.ThreadPoolExecutor); asyncio with an async SSH library is the alternative. Then the parts people forget: a bounded pool (20-50 workers) so you do not become a login storm against the devices or a TACACS server, per-connection timeouts, retries with backoff, per-device try/except so one failure does not kill the run, output normalisation across vendors (TextFSM or napalm getters), and writing results as JSON keyed by hostname. Say the complexity: O(d) connections, O(d · r) routes parsed for d devices and r routes each, time roughly d / workers times the per-device latency.
import json, concurrent.futures as cf
from netmiko import ConnectHandler
def fetch(dev):
try:
with ConnectHandler(**dev, conn_timeout=10, read_timeout=30) as c:
host = c.find_prompt().strip('#>')
routes = c.send_command('show ip route', use_textfsm=True) # list of dicts
best = [r for r in routes if r.get('protocol') != 'L']
return host, {'ok': True, 'routes': best}
except Exception as e:
return dev['host'], {'ok': False, 'error': str(e)}
def run(devices, workers=25):
with cf.ThreadPoolExecutor(max_workers=workers) as ex:
return dict(ex.map(fetch, devices))
if __name__ == '__main__':
inventory = json.load(open('devices.json'))
print(json.dumps(run(inventory), indent=1))
Follow-ups: how do you store credentials (a vault, not the inventory file), how do you avoid overloading the devices (bounded pool, jitter), what if a device hangs (timeouts, cancel), how do you test the parser (saved fixture outputs).
2. How would you push a config change to 500 switches safely?
Generate from a template and source of truth, diff against the running config, dry-run and have the diff reviewed, canary on one switch with pre- and post-checks (BGP peers established, route count, interface errors), expand in waves, stop automatically if checks fail, roll back with commit-confirm or a snapshot. Name napalm'sload_merge_candidate, compare_config, commit_config and rollback as the primitives.3. What is idempotency and why does it matter for network automation?
Running the operation again produces the same result and no new change. It matters because jobs get retried after partial failure, and a non-idempotent job (append a line, add a neighbour) would duplicate configuration or flap sessions on the retry. You achieve it by declaring desired state and computing a diff rather than issuing imperative commands.4. How would you monitor BGP health across a fleet?
Collect per-peer state, prefix counts and session uptime through gNMI streaming or an API, with SNMP or CLI polling as fallback; store it as time series; alert on state changes and on prefix counts outside a learned range; dashboard the fleet as a whole. Correlate with syslog NOTIFICATION messages. Alert on impact (traffic shifted, prefixes lost), not on every flap.5. You have a parser that works on one vendor's output. A second vendor is added. What now?
Normalise at the edge: one function per vendor that returns the same dict schema, and everything downstream consumes the schema. Prefer structured output (JSON from the device, NETCONF, gNMI) over text. Keep saved fixture outputs for both vendors and run the parsers in CI so a firmware upgrade that changes the text breaks a test, not production.Behavioral round
The loop adds a behavioral interview, and even the screen opens with "tell me about yourself". One candidate's loop asked "why Meta" and "conflicts with teammates" back to back. Meta's guide says: be yourself, be open about successes and ways you have improved, and call out how you specifically added value to your team.
STAR, kept to two minutes: Situation (two sentences of context), Task (what you were responsible for), Action (what you did, in first person singular, the longest part), Result (what changed, with a number if you have one, and what you learned). Say "I" for your actions and "we" for the team's outcome.
The story bank
Prepare five stories from your internship on the network simulation, Linux infrastructure and virtual networking team, and from coursework or projects. Use real events. If a story does not exist, do not invent one; pick a smaller real one. Interviewers probe details, and invented details collapse under the second follow-up.
| Story | What it has to contain |
|---|---|
| A troubleshooting story | The symptom, the layers you checked and in what order, the tool that found it, the fix, and how you confirmed it. This doubles as a technical answer. |
| An automation or project story | The manual thing that hurt, what you built, how you tested it, who used it, what it saved. Mention a design choice you would make differently now. |
| A conflict or feedback story | A real disagreement about a technical choice or about feedback you received. How you listened, what you changed, how the relationship ended. No villains. |
| A failure and what you learned | Something you broke, missed or shipped late. Own it without hedging, then the specific habit you changed because of it. |
| Working across teams | A time you depended on or helped another team: how you communicated, what you clarified, what the hand-off looked like. |
Ten likely behavioral questions
| Question | A strong answer contains |
|---|---|
| 1. Tell me about yourself. | Ninety seconds: where you are in your degree, the internship in one line, the two protocols you know best, why this role. End by stopping. |
| 2. Tell me about the hardest technical problem you solved. | The troubleshooting story. Structure visible: symptom, hypotheses, elimination, fix, verification. |
| 3. Tell me about a time you automated something. | The project story. Why manual was wrong, what you built, how you kept it safe, who benefited. |
| 4. Tell me about a conflict with a teammate. | A specific disagreement, your effort to understand their view, the resolution, what you took from it. Not "we never had conflicts". |
| 5. Tell me about a time you failed. | A real failure with real consequences, owned in the first sentence. The learning is concrete and you can show you applied it later. |
| 6. Tell me about feedback that was hard to hear. | The feedback quoted roughly, your first reaction, what you did within a week, evidence it stuck. |
| 7. Tell me about a time you had to learn something fast. | What you did not know, how you chose what to learn first, who you asked, what you shipped. |
| 8. Tell me about working with people outside your team. | The cross-team story. Clear interfaces, written communication, a hand-off that worked. |
| 9. How do you prioritise when everything is urgent? | Impact and blast radius first, then reversibility, then effort. An example where you said no or deferred something, and told people. |
| 10. What would your teammates say about you? | Two true traits with a one-line example each. One of them should be about how you handle being wrong. |
"Why Meta?" and "Why this role?"
The guide tells you to prepare both. A good answer is specific to the network and to the NPE blend of software and networking. Build it from these blocks and say it in under a minute.
| Block | What to say |
|---|---|
| Scale | A global backbone, dozens of data centers, fabrics with thousands of switches, serving billions of people. Problems that only exist at that size. |
| BGP in the data center | Meta published why it runs eBGP as the fabric protocol: private ASN per tier, summarisation, a small policy set, an in-house BGP agent tested like software. You find that design interesting and can say why. |
| FBOSS and open hardware | Switch software as Linux applications, open hardware through OCP (Wedge, Minipack, now Minipack3 and the disaggregated scheduled fabric). The network is built, not bought. |
| AI fabrics | RoCE backend networks for training clusters, where ECMP, congestion control and failure detection have to be re-thought. The newest and hardest networking problems are here. |
| The NPE blend | You want to write software that runs the network, not either one alone. Your internship on simulation, Linux infrastructure and virtual networking is the same blend at smaller scale. |
Put together: "I want to work where the network is treated as software. Meta runs eBGP inside its data centers with its own BGP agent, runs FBOSS on open hardware, and is building RoCE fabrics for AI training. Those are the problems I want to learn on, and the NPE role is the one where networking knowledge and Python both get used every day. My internship on a virtual networking and Linux infrastructure team was a small version of that, and I want the large version."
Weak answers sound like: "Meta is a big company with great benefits." "I use Instagram every day." "It would look good on my resume." "I want to learn a lot." Anything you could say to any other company. Anything about compensation. Anything that mentions no technology.
Six good questions to ask
- What does an intern on this team typically own by the end of the summer, and how is that scoped?
- How does a config change reach production here: what does the pipeline between source of truth and the switch look like, and where do humans review?
- What is the split between operational work and building tools on your team in a typical week?
- What was the last incident your team learned the most from, and what changed afterwards?
- How do backbone, data center and AI fabric teams work together when a problem crosses boundaries?
- What distinguishes interns who get return offers from those who do not?
Network design round preview (loop only, low priority until the screen is passed)
Rotational and new-grad loops include a network design round in place of system design. The reported question: design an enterprise LAN for 100 hosts that scales to 1000, covering topology, STP, routing, security, management and bandwidth. Talk through it in this order and draw while you talk.
- Requirements. Host types (users, servers, printers, Wi-Fi, IoT), uplink to the Internet, availability target, growth to 1000 in how long.
- Topology. 100 hosts: collapsed core, two core switches, access switches per closet, dual uplinks. 1000 hosts: three tiers (access, distribution, core) or a leaf-spine; distribution pairs per building or floor group.
- Layer 2. VLANs per function (users, servers, voice, guest, management). Choose: classic L2 access with RSTP (or MSTP) and root bridge pinned at the distribution with priorities set, or routed access so each access switch is an L3 hop and STP shrinks to a single switch. Say why: routed access removes L2 loops and converges on routing timers. Port security, BPDU guard and storm control on edge ports.
- Addressing. One /24 per VLAN with room to grow, a summarisable block per distribution pair, DHCP with relay, DNS internal and forwarding.
- Routing. OSPF as the IGP (one area at 100 hosts, areas per distribution block at 1000), default route from the edge, static or BGP to the ISP. First-hop redundancy with VRRP or an MLAG pair at the distribution.
- Redundancy. Dual uplinks everywhere, LAG where both ends support it, two cores, two edge routers, two ISPs if budget allows.
- Security. Segmentation by VLAN with inter-VLAN ACLs or a firewall between zones, guest isolated, 802.1X on access ports, management VLAN reachable only from a jump host.
- Management. Out-of-band management network, centralised syslog, SNMP or streaming telemetry, NTP, config backup, and a source of truth plus templates from day one so going from 100 to 1000 is a loop, not a project.
- Bandwidth. 1G to hosts, 10G access uplinks, 40G or 100G between distribution and core; oversubscription ratios stated (20:1 at access is fine for users, 4:1 or lower for servers). QoS for voice.
- Growth. What changes at 1000: more distribution pairs, OSPF areas, summarisation, perhaps a dedicated server block or small leaf-spine, and the automation pays off.
About the multiple-choice OA
Several non-intern candidates described a short online assessment before the screens: around 15-20 multiple-choice questions in 16-20 minutes covering TCP/IP, BGP, Python and troubleshooting. The official intern guide does not mention one, so do not count on it. A 15-minute warm-up is cheap insurance: TCP header fields and the handshake, BGP states and the best-path order, Python list and dict operations, basic complexity. If it appears, read every option before answering and keep a steady pace; a minute a question is the budget.
Interview questions
1. Tell me about your networking background and which protocols you know best.
Ninety seconds. Degree and focus, the internship in one sentence, then "BGP and TCP in depth, OSPF well enough to compare". Stop and let them pick.2. Pick a protocol and explain it in depth.
Pick BGP. Run the five blocks: overview, session and messages and best path, three failure modes with how you would see them, trade-offs versus OSPF, where you used it. Pause between blocks.3. Compare BGP and OSPF. When would you choose each?
Path-vector over TCP with policy and topology hiding versus link-state over IP with fast convergence and full trust. OSPF inside one administrative domain; BGP between domains or at hyperscale where per-link policy and failure isolation matter, which is Meta's data center case.4. Walk me through the TCP three-way handshake and what each header field does.
SYN with initial sequence number and options (MSS, window scale, SACK permitted), SYN-ACK acknowledging seq+1 with the server's ISN, ACK. Then ports, seq, ack, flags, window, checksum, options, one sentence each. Follow-up to expect: what if the SYN-ACK never arrives (retransmit with backoff; filtered versus closed port).5. A BGP peer is stuck in Active. What do you check?
Active means TCP cannot complete. Reachability to the peer address, port 179 blocked, eBGP multihop missing, wrong update source, peer not configured for me. A wrong remote AS would show OpenSent then a NOTIFICATION instead.6. How would you automate collecting the routing table from 1000 devices?
Inventory, bounded thread pool or asyncio, timeouts, retries, per-device error handling, structured parsing, JSON output keyed by hostname, credentials from a vault. Prefer an API over CLI where one exists.7. Tell me about a time you troubleshot something hard.
STAR with the layers visible: symptom, ordered hypotheses, the tool that found it, the fix, the verification. Two minutes. A real event.8. Why Meta, and why this role?
Scale, eBGP in the data center with an in-house agent, FBOSS on open hardware, AI fabrics, and the NPE blend of software and networking that matches your internship. Under a minute, no generic lines.9. You do not seem sure about that. Do you know how X works?
"I do not know X well. Here is what I do know, and here is how I would reason about it." Then reason aloud from the layer and the problem it solves. Take any hint immediately.10. Do you have questions for me?
Two prepared from the list above. One about the team's day to day, one technical about how changes reach production. Listen to the answer and ask one follow-up.Traps
- Reciting without structure. A list of BGP attributes in random order sounds like memorisation; the same facts in overview, mechanism, failure, trade-off order sound like understanding.
- Bluffing. Interviewers follow up twice on anything that sounds shaky. "I do not know" costs one question; a wrong confident answer costs the round.
- Claiming depth you do not have. If you say "I know OSPF in depth", you will spend 15 minutes on LSA types. Say what you actually know.
- Stories longer than two minutes. Practice with a timer. The interviewer will ask for the details they want.
- Badmouthing a previous team, manager, school or company. Conflict stories are about how you behaved, not about who was wrong.
- No questions prepared, or questions that Google could answer.
- Inventing details about your internship. Use what you did; describe the parts you only observed as observed.
- Ignoring hints or arguing with them. The guide says to take hints and stay open to other solutions; that is scored.
- Generic "Why Meta" with no technology in it.
Scenario
Interview minute 32. "Users in one office report that an internal web application is slow and sometimes times out. Other applications are fine. Walk me through how you would investigate." You have eight minutes and no tools; the interviewer answers questions as if they were the network.
Expected reasoning
- Scope first, out loud: one office, one application, intermittent. That points at the path between that office and that application, or at that application's handling of that office's traffic, not at the core. Ask: when did it start, did anything change, how many users, is it slow or failing.
- Reproduce from a host in the office:
pingandtracerouteto the server (loss, latency, which hop),curl -wfor connect time versus first-byte time (network versus application),digto rule out slow resolution. - If connect time is high or there is loss: capture on the client (
tcpdump host server) and look for retransmissions, duplicate ACKs, zero windows. Loss on the path shows as retransmissions; a slow application shows a fast handshake and a long wait for data. - If the path is suspect: check interface error counters and utilisation along the traceroute path, look for a congested office uplink or a duplex or MTU problem (large responses hang, small ones work). Compare with a working office to isolate what differs.
- If the application is suspect: compare response times from a host next to the server; check whether the load balancer sends this office to a different backend.
- State the fix for the cause you found, how you would verify it (the same measurement before and after), and what monitoring would have caught it earlier. Say which step you would skip if the interviewer tells you a change happened yesterday.
The last 48 hours
- Sleep. Two short review sessions beat one long night. Reread the BGP best-path order and the TCP header once each morning and stop.
- Say the three scripts aloud with a timer: BGP five blocks, TCP five blocks, OSPF 45 seconds. Record one run and listen for filler.
- Say the five STAR stories aloud once each, two minutes maximum. Say "Why Meta" once.
- Set up the machine per the guide: Zoom installed and updated, screen sharing enabled and tested, webcam, a headset or headphones with a mic, no virtual or blurred background and no filters, a quiet room and a reliable connection. Have a phone hotspot as a backup.
- Open CoderPad's sandbox once and type for ten minutes so the editor is familiar; no code execution, so dry-run habits matter.
- Have two questions for the interviewer written on paper next to you, plus the five-block script as five words.
- Water on the desk, notifications off, the job description read once more, the name of the role you are interviewing for ready to say.
- Walk in planning to say "I do not know that one, but here is how I would reason about it" at least once. It will come out naturally when needed.
← 20 · The 45-minute coding interview script · all topics · review sheet →