15 · JSON, CSV, scripts and the automation mindset Code
json/csv modules, argparse, generators, exceptions, idempotent scripts, bounded retries
Why it matters for NPE. The posting says automation is the key to meeting Meta's demands. Clean scripting and error handling are evaluated directly in the coding round and discussed in the network round.
Primer: a script is a parser with an exit code
Topic 5 covered reading files, splitting lines and aggregating into dictionaries. This page is about what wraps that loop in real life: structured formats (JSON, CSV out as well as in), a command-line shell, logging, exceptions that mean something, and the discipline that lets the same script run against one device and then a thousand. The posting says automation is how the team keeps up. The interviewer tests it two ways: in the coding round through the shape of your code (small functions, specific exceptions, streaming), and in the network round with the reported question "write a script that logs into 100 devices, runs sh ip route, parses hostnames and best routes. Now scale it to 1000."
Watch
0:40 to 10:59 is the part you need: loads and load, looping records, deleting keys, dumps with indent, dump to a file. The Yahoo Finance example from 11:45 is optional.
5:59 where to catch and which types, 12:43 a context manager for cleanup, 16:12 a retry decorator, 18:37 when exceptions are the wrong tool.
3:49 file processing, 8:05 generator pipelines. The streaming shape every script on this page uses.
Skip to the concurrent.futures part (about 22:00 on): ThreadPoolExecutor, submit, as_completed, map. That is the 100-devices answer.
Only if try/except/else/finally is not automatic yet. Ten minutes.
JSON: config files, API payloads, JSON Lines
import json
data = json.load(f) # file object, whole document
data = json.loads(text) # string
json.dump(data, f, indent=2) # write to file
text = json.dumps(data, indent=2, sort_keys=True) # to string, stable key order
# nested access without KeyError: chain .get with empty-dict defaults
mtu = cfg.get('interfaces', {}).get('eth0', {}).get('mtu', 1500)
def dig(d, *keys, default=None): # same thing, reusable
for k in keys:
if not isinstance(d, dict) or k not in d:
return default
d = d[k]
return d
# JSON Lines: one object per line, stream it
with open('events.jsonl', encoding='utf-8') as f:
for line_no, line in enumerate(f, 1):
line = line.strip()
if not line:
continue
try:
rec = json.loads(line)
except json.JSONDecodeError as e: # a ValueError subclass
log.warning('line %d: bad json: %s', line_no, e)
continue
handle(rec)
# dates and other non-JSON types: default= is called for anything dumps cannot encode
json.dumps({'when': datetime.now(), 'ip': ipaddress.ip_address('10.0.0.1')}, default=str)
| Flag | Effect | Use it when |
|---|---|---|
indent=2 | pretty output, one key per line | humans or diff will read it |
sort_keys=True | keys in alphabetical order | comparing two configs: dumps both and difflib.unified_diff the lines |
default=str | anything unencodable becomes its str() | datetimes, ipaddress objects, Paths, sets (as text) |
ensure_ascii=False | keep non-ASCII characters instead of \uXXXX | output is for people |
Two facts interviewers probe. JSON keys are always strings: json.loads(json.dumps({1: 'a'})) gives {'1': 'a'}. And a single large JSON array cannot be streamed with the standard library; JSON Lines can, which is why logs and exports use it. Say "I would ask for JSON Lines or use ijson" when the file is big.
CSV out: writing the report
import csv, sys
FIELDS = ['host', 'routes', 'default_via', 'bgp', 'ospf']
def write_report(rows, path=None):
out = open(path, 'w', newline='', encoding='utf-8') if path else sys.stdout
try:
w = csv.DictWriter(out, fieldnames=FIELDS, extrasaction='ignore')
w.writeheader()
for row in rows: # rows can be a generator: nothing buffered
w.writerow(row)
finally:
if path:
out.close()
newline='' is required. The csv module writes its own \r\n line endings and handles newlines inside quoted fields; without it Windows gets blank lines between rows and embedded newlines break. extrasaction='ignore' drops dict keys that are not in fieldnames instead of raising. Missing keys are written as restval (default empty string). Numbers are written as text; the reader converts back. For a report that could be huge, pass a generator and write as you go.
The script shell: argparse, stdin, logging, exit codes, pathlib
import argparse, logging, sys
from pathlib import Path
log = logging.getLogger('routes')
def parse_args(argv=None): # argv parameter makes this testable
p = argparse.ArgumentParser(description='Summarise saved "show ip route" output.')
p.add_argument('inputs', nargs='*', type=Path, help='files to parse; none means stdin')
p.add_argument('-o', '--output', type=Path, help='CSV report path (default stdout)')
p.add_argument('--workers', type=int, default=20)
p.add_argument('--dry-run', action='store_true', help='show what would run, change nothing')
p.add_argument('-v', '--verbose', action='store_true')
return p.parse_args(argv)
def main(argv=None):
args = parse_args(argv)
logging.basicConfig(stream=sys.stderr, level=logging.DEBUG if args.verbose else logging.INFO,
format='%(asctime)s %(levelname)s %(name)s: %(message)s')
texts = (p.read_text(encoding='utf-8') for p in args.inputs) if args.inputs else [sys.stdin.read()]
rows, bad = [], 0
for text in texts:
host, routes = parse_show_ip_route(text)
if host is None:
bad += 1
log.warning('no prompt found in input, skipping')
continue
rows.append(summarize(host, routes))
write_report(rows, args.output)
log.info('%d hosts summarised, %d skipped', len(rows), bad)
return 0 if bad == 0 else 3 # 0 ok, 3 partial failure, argparse uses 2 for usage
if __name__ == '__main__':
sys.exit(main())
| Piece | Rule |
|---|---|
argparse | nargs='*' for optional file lists, type=Path and type=int so validation is free, action='store_true' for flags. -h is generated. Bad usage exits 2 with a message, which is the right behaviour. |
| stdin | Line-oriented input: for line in sys.stdin: streams. sys.stdin.read() loads everything and is fine only for one small document, like one device's output piped in. |
logging not print | Logs go to stderr with a level and a timestamp; the report goes to stdout. The two never mix, so script.py | sort still works. log.warning('%s failed', host) formats lazily; log.exception(...) inside an except adds the traceback. |
| exit codes | 0 success, 1 unexpected failure, 2 usage, your own small integers for partial results. Cron and CI read them. return from main and sys.exit(main()) once, at the bottom. |
pathlib | Path(p).read_text(), .exists(), .suffix, Path('out') / f'{host}.txt', .mkdir(parents=True, exist_ok=True), .glob('*.log'). Fewer string joins to get wrong. |
Small functions, pure core, and a self-test you can run without devices
Split every script into three layers. Parse: text in, data out, no I/O, no printing. Decide: data in, data out (summaries, diffs, the list of changes to make). Act: the only layer that touches files, sockets or stdout. The first two layers are pure, so they are testable with assert and reusable from a notebook, a cron job or another script. Returning data instead of printing it is the single most visible difference between a script and a tool.
def _selftest():
sample = (
'R1#show ip route\n'
'S* 0.0.0.0/0 [1/0] via 10.0.0.1\n'
'C 10.0.0.0/30 is directly connected, GigabitEthernet0/0\n'
'O 10.1.1.0/24 [110/2] via 10.0.0.1, 00:12:33, GigabitEthernet0/0\n'
'B 192.0.2.0/24 [20/0] via 203.0.113.1, 1d02h\n'
)
host, routes = parse_show_ip_route(sample)
assert host == 'R1'
assert [r['proto'] for r in routes] == ['S*', 'C', 'O', 'B']
assert routes[0]['vias'] == ['10.0.0.1'] and routes[1]['vias'] == []
assert routes[1]['iface'] == 'GigabitEthernet0/0'
assert summarize(host, routes) == {'host': 'R1', 'routes': 4, 'default_via': '10.0.0.1', 'bgp': 1, 'ospf': 1}
assert parse_show_ip_route('')[0] is None
print('selftest ok')
Ten lines, runs in a millisecond, and in CoderPad (where nothing executes) it doubles as your worked example when you trace the code by hand. Wire it to --selftest or run it under if __name__ == '__main__' in a test file.
Parsing device output: show ip route into dicts
Sample of what the device returns. The prompt line carries the hostname. Each installed (best) route is one line beginning with a protocol code; equal-cost extra next hops are continuation lines that start with whitespace; summary lines like "is variably subnetted" are noise.
import re
from collections import Counter
PROMPT_RE = re.compile(r'^(?P<host>[\w.-]+)[#>]')
ROUTE_RE = re.compile(
r'^(?P<proto>[A-Z]\*?(?: [A-Z0-9]{1,2})?)\s+' # S*, C, O, B, O IA, D EX
r'(?P<prefix>\d+\.\d+\.\d+\.\d+/\d+)\s+'
r'(?:\[(?P<ad>\d+)/(?P<metric>\d+)\] via (?P<via>\d+\.\d+\.\d+\.\d+)'
r'|is directly connected, (?P<iface>\S+))')
ECMP_RE = re.compile(r'^\s+\[(?P<ad>\d+)/(?P<metric>\d+)\] via (?P<via>\d+\.\d+\.\d+\.\d+)')
def parse_show_ip_route(text):
"""Return (hostname or None, list of route dicts). Pure: no I/O."""
host, routes = None, []
for line in text.splitlines():
if host is None:
m = PROMPT_RE.match(line)
if m:
host = m.group('host')
continue
m = ROUTE_RE.match(line)
if m:
d = m.groupdict()
d['ad'] = int(d['ad']) if d['ad'] else 0
d['metric'] = int(d['metric']) if d['metric'] else 0
d['vias'] = [d.pop('via')] if d['via'] else []
routes.append(d)
continue
m = ECMP_RE.match(line)
if m and routes:
routes[-1]['vias'].append(m.group('via')) # extra equal-cost next hop
return host, routes
def summarize(host, routes):
by_proto = Counter(r['proto'][0] for r in routes) # 'O IA' and 'O' both count as OSPF
default = next((r for r in routes if r['prefix'] == '0.0.0.0/0'), None)
return {'host': host, 'routes': len(routes),
'default_via': default['vias'][0] if default and default['vias'] else None,
'bgp': by_proto['B'], 'ospf': by_proto['O']}
Walk the sample: the prompt gives R1. Line S* matches with ad=1, metric=0, vias=['10.0.0.1']. The "variably subnetted" line matches nothing and is skipped. C and L take the "directly connected" branch, so iface is set and vias is empty. O IA matches because the protocol group allows a one-space suffix of one or two letters; its continuation line matches ECMP_RE and appends 10.0.0.5 to the previous route. Output: 6 routes, default via 10.0.0.1, 1 BGP, 2 OSPF. O(L) in the number of lines, O(R) memory for R routes.
Say the limits out loud: this regex is IOS-shaped. NX-OS, Junos and Arista print differently, so in production you would use the vendor's structured output (| json, NETCONF, gNMI) or a TextFSM template and keep the regex as a fallback. Compile regexes once at module level, name the groups, and keep the per-line work to one match.
Exceptions that carry meaning
class DeviceError(Exception):
"""A device could not be reached or returned unusable output."""
def fetch_routes(host):
try:
text = run_command(host, 'show ip route')
except (TimeoutError, ConnectionError, OSError) as e:
raise DeviceError(f'{host}: unreachable') from e # keeps the original as __cause__
name, routes = parse_show_ip_route(text)
if name is None:
raise DeviceError(f'{host}: no prompt in output, got {text[:60]!r}')
return name, routes
- Catch specific types.
except (ValueError, KeyError)around the conversion that can fail, not around the whole loop. A bareexcept:swallowsKeyboardInterruptandSystemExit;except Exceptionis acceptable only at the outermost boundary of a worker, where you record the failure and move on. raise NewError(...) from etranslates a library error into your domain error and keeps the traceback chain.from Nonehides the cause on purpose.- One custom exception per script is usually enough. It lets the caller write
except DeviceErrorand let real bugs (aTypeErrorin your parser) surface instead of being counted as "device failed". try/except/else/finally:elseruns when nothing was raised,finallyalways. Use awithblock instead offinallyfor closing things.- Decide the policy and say it: skip and count, skip and record, or fail fast. For reports, skip and count. For a change script, fail fast before any change is made and never half way through one.
Retries, timeouts, idempotence, dry runs, canaries
import random, time
def with_retries(fn, *args, attempts=3, base=0.5, cap=8.0, retry_on=(TimeoutError, ConnectionError)):
for attempt in range(1, attempts + 1):
try:
return fn(*args)
except retry_on as e:
if attempt == attempts:
raise DeviceError(f'{args[0]}: gave up after {attempts} attempts') from e
delay = min(cap, base * 2 ** (attempt - 1)) * random.uniform(0.5, 1.5) # backoff + jitter
log.warning('%s attempt %d failed (%s); retrying in %.1fs', args[0], attempt, e, delay)
time.sleep(delay)
| Rule | Why | What it looks like |
|---|---|---|
| Timeout on every network call | A hung SSH session with no timeout holds a worker forever; with 20 workers and 20 hung devices the batch never finishes and nothing reports it. | socket.create_connection(addr, timeout=10); netmiko conn_timeout, read_timeout; requests.get(url, timeout=(3, 10)). Never timeout=None. |
| Bounded retries with exponential backoff and jitter | Retrying instantly turns one failure into a flood; jitter stops 1000 clients retrying in lockstep. Only retry errors that can succeed later (timeouts, resets), never authentication failures or parse errors. | 3 attempts, 0.5 s, 1 s, 2 s, each multiplied by a random factor; log each attempt. |
| Idempotent operations | Running the script twice must leave the device in the same state as running it once. Then a retry after a half-finished run is safe. | "Ensure VLAN 20 exists" (check, then add if missing), not "add VLAN 20". Write the whole file to a temp path and os.replace it, rather than appending. |
| Dry run | Shows the plan without acting. Every change script needs --dry-run and it should be the default. | Decide layer returns the list of changes; act layer prints them when args.dry_run and applies them otherwise. |
| Canary | One device, verify, then a small batch, then all. A bug that would have broken 1000 devices breaks one. | --limit 1 first; check the device; then --limit 10; then the fleet. Stop the batch when the failure rate crosses a threshold. |
Generators, itertools and context managers
def read_lines(paths): # one stream over many files
for p in paths:
with open(p, encoding='utf-8') as f:
yield from f
def parse(lines): # text -> dicts, lazily
for line in lines:
rec = parse_line(line)
if rec is not None:
yield rec
def only_down(records):
return (r for r in records if r['status'] == 'DOWN') # generator expression
pipeline = only_down(parse(read_lines(args.inputs))) # nothing has run yet
for rec in pipeline: # one record in memory at a time
...
from itertools import islice, chain, groupby
first_ten = list(islice(pipeline, 10)) # take 10 without consuming the rest
everything = chain(read_lines(a_files), read_lines(b_files))
for host, group in groupby(sorted(recs, key=lambda r: r['host']), key=lambda r: r['host']):
print(host, sum(1 for _ in group)) # groupby only groups ADJACENT equal keys: sort first
from contextlib import contextmanager
import time
@contextmanager
def timed(label):
start = time.perf_counter()
try:
yield
finally:
log.info('%s took %.2fs', label, time.perf_counter() - start)
with timed('fleet audit'):
rows, failures = collect(hosts, 'show ip route')
A generator is a function that pauses at yield and resumes on the next request. Chained generators form a pipeline where every stage holds one item, so memory is O(1) in the number of lines no matter how many files you chain. Two things break the pipeline on purpose: sorting (needs everything) and groupby on unsorted input (gives wrong groups, not an error). yield from delegates to another iterator. A context manager pairs setup with guaranteed cleanup; use one for anything with a close, disconnect or unlock step, including the SSH session in the worker below.
100 devices, then 1000: concurrency without losing control
The question. Say the serial version first: a loop over hosts, each login taking about two seconds, so 100 devices is a few minutes and 1000 devices is half an hour or more. Then: the work is almost entirely waiting on the network, so run the logins in parallel with a bounded thread pool, keep every call bounded by a timeout, collect failures without stopping the batch, and write results as they arrive.
import logging, random, time
from concurrent.futures import ThreadPoolExecutor, as_completed
log = logging.getLogger('fleet')
def run_command(host, command, timeout=15):
"""Stand-in for the SSH call. Production: netmiko ConnectHandler(host=..., conn_timeout=timeout)
.send_command(command, read_timeout=timeout), or paramiko / napalm / scrapli. The real call
MUST take a timeout; a thread cannot be killed from outside."""
time.sleep(random.uniform(0.1, 0.5)) # pretend to wait on the network
if random.random() < 0.05:
raise TimeoutError(f'{host}: no banner within {timeout}s')
return f'{host}#show ip route\nS* 0.0.0.0/0 [1/0] via 10.0.0.1\n'
def collect(hosts, command, workers=20, batch_timeout=600):
"""Run command on every host. Returns ({host: output}, {host: error text}). Never raises for one host."""
results, failures = {}, {}
with ThreadPoolExecutor(max_workers=workers) as pool:
futures = {pool.submit(with_retries, run_command, h, command): h for h in hosts}
try:
for fut in as_completed(futures, timeout=batch_timeout):
host = futures[fut]
try:
results[host] = fut.result()
except DeviceError as e:
failures[host] = str(e)
except Exception as e: # a bug in our code, not the device
failures[host] = f'bug: {type(e).__name__}: {e}'
log.exception('%s: unexpected error', host)
except TimeoutError:
for fut, host in futures.items():
if not fut.done():
failures[host] = 'batch timeout'
return results, failures
def audit(hosts, out_path, workers=20):
outputs, failures = collect(hosts, 'show ip route', workers=workers)
rows = (summarize(*parse_show_ip_route(text)) for text in outputs.values())
write_report(rows, out_path)
for host, why in sorted(failures.items()):
log.error('FAILED %s: %s', host, why)
return len(outputs), len(failures)
| Concern at 1000 | Answer |
|---|---|
| How many threads? | Not 1000. A bounded pool of 20 to 50. Each worker holds an SSH session; the limits are the AAA/TACACS server's login rate, the control-plane CPU on small devices, and your own file descriptors. max_workers is the knob; start low and raise it while watching failure rate. |
| Rate limiting | The pool size already caps concurrency. For logins per second, sleep between submit calls or use a token bucket; for per-site limits, a threading.Semaphore per site acquired inside the worker. |
| One device hangs | The timeout inside run_command raises, with_retries tries again with backoff, then raises DeviceError, which lands in failures. The other 999 continue. The batch-level as_completed(timeout=) is the last line of defence. |
| Memory | 1000 route tables of a few hundred KB each fits, but do not rely on it. Parse in the worker and return the summary dict instead of the raw text, or write each output to Path(out) / f'{host}.txt' as it completes. |
| Partial results | Always produce the report for the hosts that worked plus a failure list with reasons. Exit code 3 for partial. Re-run only the failures: --hosts-from failures.txt. That re-run must be idempotent. |
| Inventory | Hosts come from a file or a source of truth (NetBox), never typed into the script. Validate the list before the first login. |
| Beyond a few thousand | A framework that already does inventory, pools, retries and per-host results: nornir with netmiko or napalm, or an async stack (scrapli, asyncssh). Or split the fleet by site and run from a worker near each site. |
The GIL in two sentences. CPython lets one thread execute Python bytecode at a time, so threads do not speed up pure Python computation. Waiting on a socket releases the GIL, so for SSH sessions threads give nearly linear speedup; if the bottleneck moved to CPU-heavy parsing you would use ProcessPoolExecutor for that stage.
asyncio in one paragraph. The same shape with one thread and an event loop: async def fetch(host) using an async SSH library (asyncssh, scrapli-asyncssh), asyncio.Semaphore(50) to bound concurrency, asyncio.wait_for(coro, timeout=15) per call, asyncio.gather(*tasks, return_exceptions=True) to collect results and exceptions together. It scales to more concurrent sessions than threads with less memory, but every library in the path must be async. In an interview, write the thread pool; mention asyncio as the next step.
Data structure design problems
The four design problems linked below are "automation mindset" in miniature: pick the containers that make each operation O(1) or O(log n), and state the guarantees out loud.
import random
from bisect import bisect_right
class RandomizedSet: # LC 380: list for O(1) random, dict for O(1) index lookup
def __init__(self):
self.vals, self.pos = [], {}
def insert(self, v):
if v in self.pos:
return False
self.pos[v] = len(self.vals); self.vals.append(v)
return True
def remove(self, v):
if v not in self.pos:
return False
i, last = self.pos[v], self.vals[-1]
self.vals[i], self.pos[last] = last, i # move the last element into the hole
self.vals.pop(); del self.pos[v] # then pop: O(1), no shifting
return True
def getRandom(self):
return random.choice(self.vals)
class TimeMap: # LC 981: timestamps per key arrive increasing, so bisect
def __init__(self):
self.ts, self.vals = {}, {} # key -> [timestamps], key -> [values]
def set(self, key, value, t):
self.ts.setdefault(key, []).append(t)
self.vals.setdefault(key, []).append(value)
def get(self, key, t):
if key not in self.ts:
return ''
i = bisect_right(self.ts[key], t) # number of timestamps <= t
return self.vals[key][i - 1] if i else '' # O(log k) for k versions of the key
| Problem | Containers | Guarantee to state |
|---|---|---|
| Insert Delete GetRandom O(1) | list + dict of value to index; delete by swap with last and pop | all three O(1); random.choice needs a list, a dict cannot be indexed |
| Time Based Key-Value Store | per key two parallel lists (or a list of tuples); bisect_right on timestamps | set O(1) amortised, get O(log k); relies on timestamps increasing, say so. Versioned config lookup: "what was this value at time t" |
| Design Authentication Manager | dict token to expiry time; renew checks expiry first; countUnexpiredTokens scans or purges lazily | generate/renew O(1); count O(n) unless you keep an OrderedDict in expiry order and pop from the front. Tokens with TTL is a session cache |
| Design Underground System | dict id to (station, time) for open journeys; dict (start, end) to [total, count] | all O(1); pop the check-in on check-out so memory is bounded by open journeys. Pairing start and end events is the flow-log pattern |
Interview questions
1. json.dump raises TypeError: Object of type datetime is not JSON serializable. Fix?
Pass default=str (or a function that formats the types you care about) so unencodable objects are converted instead of raising. Say that JSON has no date type, so the reader must parse the string back.2. How do you make two JSON configs comparable?
Load both,dumps each with sort_keys=True, indent=2, split into lines and feed difflib.unified_diff. Sorting keys removes ordering noise; indent puts one key per line so the diff is readable. Follow-up: lists are order-sensitive, so sort list members too when order does not matter.3. Why newline='' when opening a file for the csv module?
The csv module writes its own line endings and handles newlines inside quoted fields. If the file object also translates newlines you get \r\r\n on Windows (blank lines) and corrupted quoted fields. Same flag for reading.4. How do you structure a script so it can be tested?
Pure parse and decide functions that take data and return data, a thinmain(argv) that does I/O, parse_args(argv=None) so tests can pass a list, and sys.exit(main()) under if __name__ == '__main__'. Then a self-test or pytest file calls the pure functions with a small sample and asserts.5. Write a script that logs into 100 devices, runs sh ip route and parses hostnames and best routes. Now 1000.
Serial loop first: each login about two seconds, so minutes for 100 and too long for 1000. The work is network waiting, so a ThreadPoolExecutor with 20 to 50 workers, each call with a connect and read timeout, bounded retries with backoff, failures collected per host without stopping the batch, results written as they complete, a CSV report plus a failure list, exit code for partial. Hostname from the prompt line, routes with one compiled regex per line. For more than a few thousand, nornir or an async stack, and split by site.6. Python has the GIL. Why do threads help here at all?
The GIL serialises Python bytecode, but blocking I/O releases it. An SSH session spends almost all its time waiting on the socket, so 20 threads give close to 20x. If parsing became CPU-bound I would move that stage to a process pool.7. What happens if one device hangs and you have no timeout?
That worker never returns. The pool keeps running the others, butas_completed never finishes and the with block waits forever at shutdown, so the report is never written. Threads cannot be killed, so the timeout has to be inside the call. I also put a batch-level timeout on as_completed as a backstop.8. How do you retry safely?
Only retry errors that can succeed later: timeouts, connection resets, a 503. Never retry authentication failures or parse errors. Bounded attempts, exponential backoff with a cap, random jitter so a thousand clients do not retry in lockstep, log every attempt, and make the operation idempotent so a retry after a half-finished attempt is harmless.9. What does idempotent mean for a network change script?
Running it twice leaves the device in the same state as running it once. "Ensure the VLAN exists" rather than "add the VLAN". It is what makes retries and re-runs after a crash safe, and it is why the script should read current state, compute a diff, and apply only the diff.10. Why is a bare except: wrong, and what do you write instead?
It catches everything including KeyboardInterrupt and SystemExit, so the script cannot be stopped and bugs are hidden. Catch the specific exceptions the operation can raise, as close to the call as possible. At a worker boundary except Exception is acceptable if you log it with log.exception and record the failure. Translate library errors with raise DeviceError(...) from e.11. When does a generator pipeline not help?
When a stage needs the whole input: sorting,groupby on unsorted data, computing a median, or anything that has to see the last line before emitting the first. Then memory is O(n) at that stage no matter how lazy the rest is, and I would say so and consider external sort or partitioning.12. Insert, delete and getRandom in O(1). How?
A list holds the values sorandom.choice is O(1), and a dict maps value to its index. Delete swaps the victim with the last element, updates that element's index, and pops the tail, so no shifting. Follow-up with duplicates allowed: dict of value to a set of indices.Traps
- Bare
except:orexcept Exception: pass. Silent failures are worse than crashes; the interviewer will ask what happens on a bad line. - Unbounded threads: one thread per host. 1000 SSH sessions at once is a denial of service on the AAA server and your own machine. Bound the pool.
- No timeout on a network call. A thread cannot be killed; one hung device hangs the batch.
- Mutable default arguments:
def collect(hosts, failures=[])shares one list across calls. UseNoneand create inside. - Printing instead of returning. A function that prints cannot be tested, reused or written to CSV. Return data; print in
main. sys.stdin.read()orf.read()on input you were told is large. Iterate lines.itertools.groupbyon unsorted input: it groups adjacent runs only, so you get the same key several times with no error.- Catching the exception and retrying a non-idempotent change, so the retry applies it twice.
- Mixing report and log on stdout. Logs to stderr via
logging, data to stdout, so the output can be piped.
Do before marking this topic done
- Labs 7 to 10 below, timed, blank editor: inventory join with UNKNOWN region, most-specific prefix match with
ipaddress, merge two sorted logs as a stream, burst alerts with a per-host deque. Each one as a pure function plus amainthat reads the fixtures and prints JSON. - The 100-devices script end to end with the fake
run_command:argparse(--hosts,--workers,--dry-run,-o),collectwith a bounded pool and timeouts,with_retries,parse_show_ip_route,summarize, CSV report, failure list on stderr, exit code 3 on partial. Then say the 1000-device answer out loud in under a minute. - Write
parse_show_ip_routefrom memory against the sample above and make the ten-line self-test pass. Add a line your regex does not handle and decide: extend the regex or count it as unparsed? - JSON config diff tool: two files in,
sort_keysplusindent,difflib.unified_diffout, exit 0 when identical and 1 when different. Then make it accept JSON Lines anddefault=stra datetime field. - Insert Delete GetRandom O(1) and Time Based Key-Value Store without looking. For each, say every operation's complexity before typing.
Practice: linked problems and labs
Logged attempts feed the tracker. Open the problem in a new tab, solve in a blank editor, then log honestly.
- LC 380Insert Delete GetRandom O(1)Medium
- LC 981Time Based Key-Value StoreMedium
- LC 1797Design Authentication ManagerMedium
- LC 1396Design Underground SystemMedium
- lab 7Join inventory with trafficLab
- lab 8Match addresses to the most specific prefixLab
- lab 9Merge two sorted log filesLab
- lab 10Detect repeated failures in a time windowLab
← 14 · BGP in depth · all topics · 16 · Meta's network: Clos fabrics, ECMP, BGP in the DC, FBOSS, backbone and MPLS →