← all topics

15 · JSON, CSV, scripts and the automation mindset Code

json/csv modules, argparse, generators, exceptions, idempotent scripts, bounded retries

Why it matters for NPE. The posting says automation is the key to meeting Meta's demands. Clean scripting and error handling are evaluated directly in the coding round and discussed in the network round.

Primer: a script is a parser with an exit code

Topic 5 covered reading files, splitting lines and aggregating into dictionaries. This page is about what wraps that loop in real life: structured formats (JSON, CSV out as well as in), a command-line shell, logging, exceptions that mean something, and the discipline that lets the same script run against one device and then a thousand. The posting says automation is how the team keeps up. The interviewer tests it two ways: in the coding round through the shape of your code (small functions, specific exceptions, streaming), and in the network round with the reported question "write a script that logs into 100 devices, runs sh ip route, parses hostnames and best routes. Now scale it to 1000."

The automation mindset in one sentence: pure functions in the middle, a thin I/O shell around them, every network call bounded by a timeout, every change idempotent, and a dry run before the real run. If you say that and your code shows it, you pass this part.

Watch

Python Tutorial: Working with JSON Data using the json ModuleCorey Schafer · 20:33

0:40 to 10:59 is the part you need: loads and load, looping records, deleting keys, dumps with indent, dump to a file. The Yahoo Finance example from 11:45 is optional.

Exception Handling Tips in Python ⚠ Write Better Python Code Part 7ArjanCodes · 21:46

5:59 where to catch and which types, 12:43 a context manager for cleanup, 16:12 a retry decorator, 18:37 when exceptions are the wrong tool.

Python GeneratorsmCoding · 15:32

3:49 file processing, 8:05 generator pipelines. The streaming shape every script on this page uses.

Python Threading Tutorial: Run Code Concurrently Using the Threading ModuleCorey Schafer · 36:05

Skip to the concurrent.futures part (about 22:00 on): ThreadPoolExecutor, submit, as_completed, map. That is the 100-devices answer.

Python Tutorial: Using Try/Except Blocks for Error HandlingCorey Schafer · 10:34

Only if try/except/else/finally is not automatic yet. Ten minutes.

JSON: config files, API payloads, JSON Lines

import json

data = json.load(f)                 # file object, whole document
data = json.loads(text)             # string
json.dump(data, f, indent=2)        # write to file
text = json.dumps(data, indent=2, sort_keys=True)   # to string, stable key order

# nested access without KeyError: chain .get with empty-dict defaults
mtu = cfg.get('interfaces', {}).get('eth0', {}).get('mtu', 1500)

def dig(d, *keys, default=None):    # same thing, reusable
    for k in keys:
        if not isinstance(d, dict) or k not in d:
            return default
        d = d[k]
    return d

# JSON Lines: one object per line, stream it
with open('events.jsonl', encoding='utf-8') as f:
    for line_no, line in enumerate(f, 1):
        line = line.strip()
        if not line:
            continue
        try:
            rec = json.loads(line)
        except json.JSONDecodeError as e:         # a ValueError subclass
            log.warning('line %d: bad json: %s', line_no, e)
            continue
        handle(rec)

# dates and other non-JSON types: default= is called for anything dumps cannot encode
json.dumps({'when': datetime.now(), 'ip': ipaddress.ip_address('10.0.0.1')}, default=str)
FlagEffectUse it when
indent=2pretty output, one key per linehumans or diff will read it
sort_keys=Truekeys in alphabetical ordercomparing two configs: dumps both and difflib.unified_diff the lines
default=stranything unencodable becomes its str()datetimes, ipaddress objects, Paths, sets (as text)
ensure_ascii=Falsekeep non-ASCII characters instead of \uXXXXoutput is for people

Two facts interviewers probe. JSON keys are always strings: json.loads(json.dumps({1: 'a'})) gives {'1': 'a'}. And a single large JSON array cannot be streamed with the standard library; JSON Lines can, which is why logs and exports use it. Say "I would ask for JSON Lines or use ijson" when the file is big.

CSV out: writing the report

import csv, sys

FIELDS = ['host', 'routes', 'default_via', 'bgp', 'ospf']

def write_report(rows, path=None):
    out = open(path, 'w', newline='', encoding='utf-8') if path else sys.stdout
    try:
        w = csv.DictWriter(out, fieldnames=FIELDS, extrasaction='ignore')
        w.writeheader()
        for row in rows:                     # rows can be a generator: nothing buffered
            w.writerow(row)
    finally:
        if path:
            out.close()

newline='' is required. The csv module writes its own \r\n line endings and handles newlines inside quoted fields; without it Windows gets blank lines between rows and embedded newlines break. extrasaction='ignore' drops dict keys that are not in fieldnames instead of raising. Missing keys are written as restval (default empty string). Numbers are written as text; the reader converts back. For a report that could be huge, pass a generator and write as you go.

The script shell: argparse, stdin, logging, exit codes, pathlib

import argparse, logging, sys
from pathlib import Path

log = logging.getLogger('routes')

def parse_args(argv=None):                   # argv parameter makes this testable
    p = argparse.ArgumentParser(description='Summarise saved "show ip route" output.')
    p.add_argument('inputs', nargs='*', type=Path, help='files to parse; none means stdin')
    p.add_argument('-o', '--output', type=Path, help='CSV report path (default stdout)')
    p.add_argument('--workers', type=int, default=20)
    p.add_argument('--dry-run', action='store_true', help='show what would run, change nothing')
    p.add_argument('-v', '--verbose', action='store_true')
    return p.parse_args(argv)

def main(argv=None):
    args = parse_args(argv)
    logging.basicConfig(stream=sys.stderr, level=logging.DEBUG if args.verbose else logging.INFO,
                        format='%(asctime)s %(levelname)s %(name)s: %(message)s')
    texts = (p.read_text(encoding='utf-8') for p in args.inputs) if args.inputs else [sys.stdin.read()]
    rows, bad = [], 0
    for text in texts:
        host, routes = parse_show_ip_route(text)
        if host is None:
            bad += 1
            log.warning('no prompt found in input, skipping')
            continue
        rows.append(summarize(host, routes))
    write_report(rows, args.output)
    log.info('%d hosts summarised, %d skipped', len(rows), bad)
    return 0 if bad == 0 else 3                 # 0 ok, 3 partial failure, argparse uses 2 for usage

if __name__ == '__main__':
    sys.exit(main())
PieceRule
argparsenargs='*' for optional file lists, type=Path and type=int so validation is free, action='store_true' for flags. -h is generated. Bad usage exits 2 with a message, which is the right behaviour.
stdinLine-oriented input: for line in sys.stdin: streams. sys.stdin.read() loads everything and is fine only for one small document, like one device's output piped in.
logging not printLogs go to stderr with a level and a timestamp; the report goes to stdout. The two never mix, so script.py | sort still works. log.warning('%s failed', host) formats lazily; log.exception(...) inside an except adds the traceback.
exit codes0 success, 1 unexpected failure, 2 usage, your own small integers for partial results. Cron and CI read them. return from main and sys.exit(main()) once, at the bottom.
pathlibPath(p).read_text(), .exists(), .suffix, Path('out') / f'{host}.txt', .mkdir(parents=True, exist_ok=True), .glob('*.log'). Fewer string joins to get wrong.

Small functions, pure core, and a self-test you can run without devices

Split every script into three layers. Parse: text in, data out, no I/O, no printing. Decide: data in, data out (summaries, diffs, the list of changes to make). Act: the only layer that touches files, sockets or stdout. The first two layers are pure, so they are testable with assert and reusable from a notebook, a cron job or another script. Returning data instead of printing it is the single most visible difference between a script and a tool.

def _selftest():
    sample = (
        'R1#show ip route\n'
        'S*    0.0.0.0/0 [1/0] via 10.0.0.1\n'
        'C        10.0.0.0/30 is directly connected, GigabitEthernet0/0\n'
        'O        10.1.1.0/24 [110/2] via 10.0.0.1, 00:12:33, GigabitEthernet0/0\n'
        'B        192.0.2.0/24 [20/0] via 203.0.113.1, 1d02h\n'
    )
    host, routes = parse_show_ip_route(sample)
    assert host == 'R1'
    assert [r['proto'] for r in routes] == ['S*', 'C', 'O', 'B']
    assert routes[0]['vias'] == ['10.0.0.1'] and routes[1]['vias'] == []
    assert routes[1]['iface'] == 'GigabitEthernet0/0'
    assert summarize(host, routes) == {'host': 'R1', 'routes': 4, 'default_via': '10.0.0.1', 'bgp': 1, 'ospf': 1}
    assert parse_show_ip_route('')[0] is None
    print('selftest ok')

Ten lines, runs in a millisecond, and in CoderPad (where nothing executes) it doubles as your worked example when you trace the code by hand. Wire it to --selftest or run it under if __name__ == '__main__' in a test file.

Parsing device output: show ip route into dicts

Sample of what the device returns. The prompt line carries the hostname. Each installed (best) route is one line beginning with a protocol code; equal-cost extra next hops are continuation lines that start with whitespace; summary lines like "is variably subnetted" are noise.

R1#show ip route Gateway of last resort is 10.0.0.1 to network 0.0.0.0 S* 0.0.0.0/0 [1/0] via 10.0.0.1 10.0.0.0/8 is variably subnetted, 4 subnets, 2 masks C 10.0.0.0/30 is directly connected, GigabitEthernet0/0 L 10.0.0.2/32 is directly connected, GigabitEthernet0/0 O 10.1.1.0/24 [110/2] via 10.0.0.1, 00:12:33, GigabitEthernet0/0 O IA 10.2.0.0/16 [110/3] via 10.0.0.1, 00:12:33, GigabitEthernet0/0 [110/3] via 10.0.0.5, 00:12:33, GigabitEthernet0/1 B 192.0.2.0/24 [20/0] via 203.0.113.1, 1d02h
import re
from collections import Counter

PROMPT_RE = re.compile(r'^(?P<host>[\w.-]+)[#>]')
ROUTE_RE = re.compile(
    r'^(?P<proto>[A-Z]\*?(?: [A-Z0-9]{1,2})?)\s+'             # S*, C, O, B, O IA, D EX
    r'(?P<prefix>\d+\.\d+\.\d+\.\d+/\d+)\s+'
    r'(?:\[(?P<ad>\d+)/(?P<metric>\d+)\] via (?P<via>\d+\.\d+\.\d+\.\d+)'
    r'|is directly connected, (?P<iface>\S+))')
ECMP_RE = re.compile(r'^\s+\[(?P<ad>\d+)/(?P<metric>\d+)\] via (?P<via>\d+\.\d+\.\d+\.\d+)')

def parse_show_ip_route(text):
    """Return (hostname or None, list of route dicts). Pure: no I/O."""
    host, routes = None, []
    for line in text.splitlines():
        if host is None:
            m = PROMPT_RE.match(line)
            if m:
                host = m.group('host')
                continue
        m = ROUTE_RE.match(line)
        if m:
            d = m.groupdict()
            d['ad'] = int(d['ad']) if d['ad'] else 0
            d['metric'] = int(d['metric']) if d['metric'] else 0
            d['vias'] = [d.pop('via')] if d['via'] else []
            routes.append(d)
            continue
        m = ECMP_RE.match(line)
        if m and routes:
            routes[-1]['vias'].append(m.group('via'))     # extra equal-cost next hop
    return host, routes

def summarize(host, routes):
    by_proto = Counter(r['proto'][0] for r in routes)     # 'O IA' and 'O' both count as OSPF
    default = next((r for r in routes if r['prefix'] == '0.0.0.0/0'), None)
    return {'host': host, 'routes': len(routes),
            'default_via': default['vias'][0] if default and default['vias'] else None,
            'bgp': by_proto['B'], 'ospf': by_proto['O']}

Walk the sample: the prompt gives R1. Line S* matches with ad=1, metric=0, vias=['10.0.0.1']. The "variably subnetted" line matches nothing and is skipped. C and L take the "directly connected" branch, so iface is set and vias is empty. O IA matches because the protocol group allows a one-space suffix of one or two letters; its continuation line matches ECMP_RE and appends 10.0.0.5 to the previous route. Output: 6 routes, default via 10.0.0.1, 1 BGP, 2 OSPF. O(L) in the number of lines, O(R) memory for R routes.

Say the limits out loud: this regex is IOS-shaped. NX-OS, Junos and Arista print differently, so in production you would use the vendor's structured output (| json, NETCONF, gNMI) or a TextFSM template and keep the regex as a fallback. Compile regexes once at module level, name the groups, and keep the per-line work to one match.

Exceptions that carry meaning

class DeviceError(Exception):
    """A device could not be reached or returned unusable output."""

def fetch_routes(host):
    try:
        text = run_command(host, 'show ip route')
    except (TimeoutError, ConnectionError, OSError) as e:
        raise DeviceError(f'{host}: unreachable') from e          # keeps the original as __cause__
    name, routes = parse_show_ip_route(text)
    if name is None:
        raise DeviceError(f'{host}: no prompt in output, got {text[:60]!r}')
    return name, routes

Retries, timeouts, idempotence, dry runs, canaries

import random, time

def with_retries(fn, *args, attempts=3, base=0.5, cap=8.0, retry_on=(TimeoutError, ConnectionError)):
    for attempt in range(1, attempts + 1):
        try:
            return fn(*args)
        except retry_on as e:
            if attempt == attempts:
                raise DeviceError(f'{args[0]}: gave up after {attempts} attempts') from e
            delay = min(cap, base * 2 ** (attempt - 1)) * random.uniform(0.5, 1.5)   # backoff + jitter
            log.warning('%s attempt %d failed (%s); retrying in %.1fs', args[0], attempt, e, delay)
            time.sleep(delay)
RuleWhyWhat it looks like
Timeout on every network callA hung SSH session with no timeout holds a worker forever; with 20 workers and 20 hung devices the batch never finishes and nothing reports it.socket.create_connection(addr, timeout=10); netmiko conn_timeout, read_timeout; requests.get(url, timeout=(3, 10)). Never timeout=None.
Bounded retries with exponential backoff and jitterRetrying instantly turns one failure into a flood; jitter stops 1000 clients retrying in lockstep. Only retry errors that can succeed later (timeouts, resets), never authentication failures or parse errors.3 attempts, 0.5 s, 1 s, 2 s, each multiplied by a random factor; log each attempt.
Idempotent operationsRunning the script twice must leave the device in the same state as running it once. Then a retry after a half-finished run is safe."Ensure VLAN 20 exists" (check, then add if missing), not "add VLAN 20". Write the whole file to a temp path and os.replace it, rather than appending.
Dry runShows the plan without acting. Every change script needs --dry-run and it should be the default.Decide layer returns the list of changes; act layer prints them when args.dry_run and applies them otherwise.
CanaryOne device, verify, then a small batch, then all. A bug that would have broken 1000 devices breaks one.--limit 1 first; check the device; then --limit 10; then the fleet. Stop the batch when the failure rate crosses a threshold.

Generators, itertools and context managers

def read_lines(paths):                       # one stream over many files
    for p in paths:
        with open(p, encoding='utf-8') as f:
            yield from f

def parse(lines):                            # text -> dicts, lazily
    for line in lines:
        rec = parse_line(line)
        if rec is not None:
            yield rec

def only_down(records):
    return (r for r in records if r['status'] == 'DOWN')     # generator expression

pipeline = only_down(parse(read_lines(args.inputs)))        # nothing has run yet
for rec in pipeline:                                         # one record in memory at a time
    ...

from itertools import islice, chain, groupby

first_ten = list(islice(pipeline, 10))               # take 10 without consuming the rest
everything = chain(read_lines(a_files), read_lines(b_files))
for host, group in groupby(sorted(recs, key=lambda r: r['host']), key=lambda r: r['host']):
    print(host, sum(1 for _ in group))                # groupby only groups ADJACENT equal keys: sort first

from contextlib import contextmanager
import time

@contextmanager
def timed(label):
    start = time.perf_counter()
    try:
        yield
    finally:
        log.info('%s took %.2fs', label, time.perf_counter() - start)

with timed('fleet audit'):
    rows, failures = collect(hosts, 'show ip route')

A generator is a function that pauses at yield and resumes on the next request. Chained generators form a pipeline where every stage holds one item, so memory is O(1) in the number of lines no matter how many files you chain. Two things break the pipeline on purpose: sorting (needs everything) and groupby on unsorted input (gives wrong groups, not an error). yield from delegates to another iterator. A context manager pairs setup with guaranteed cleanup; use one for anything with a close, disconnect or unlock step, including the SSH session in the worker below.

100 devices, then 1000: concurrency without losing control

The question. Say the serial version first: a loop over hosts, each login taking about two seconds, so 100 devices is a few minutes and 1000 devices is half an hour or more. Then: the work is almost entirely waiting on the network, so run the logins in parallel with a bounded thread pool, keep every call bounded by a timeout, collect failures without stopping the batch, and write results as they arrive.

import logging, random, time
from concurrent.futures import ThreadPoolExecutor, as_completed

log = logging.getLogger('fleet')

def run_command(host, command, timeout=15):
    """Stand-in for the SSH call. Production: netmiko ConnectHandler(host=..., conn_timeout=timeout)
    .send_command(command, read_timeout=timeout), or paramiko / napalm / scrapli. The real call
    MUST take a timeout; a thread cannot be killed from outside."""
    time.sleep(random.uniform(0.1, 0.5))                       # pretend to wait on the network
    if random.random() < 0.05:
        raise TimeoutError(f'{host}: no banner within {timeout}s')
    return f'{host}#show ip route\nS*    0.0.0.0/0 [1/0] via 10.0.0.1\n'

def collect(hosts, command, workers=20, batch_timeout=600):
    """Run command on every host. Returns ({host: output}, {host: error text}). Never raises for one host."""
    results, failures = {}, {}
    with ThreadPoolExecutor(max_workers=workers) as pool:
        futures = {pool.submit(with_retries, run_command, h, command): h for h in hosts}
        try:
            for fut in as_completed(futures, timeout=batch_timeout):
                host = futures[fut]
                try:
                    results[host] = fut.result()
                except DeviceError as e:
                    failures[host] = str(e)
                except Exception as e:                      # a bug in our code, not the device
                    failures[host] = f'bug: {type(e).__name__}: {e}'
                    log.exception('%s: unexpected error', host)
        except TimeoutError:
            for fut, host in futures.items():
                if not fut.done():
                    failures[host] = 'batch timeout'
    return results, failures

def audit(hosts, out_path, workers=20):
    outputs, failures = collect(hosts, 'show ip route', workers=workers)
    rows = (summarize(*parse_show_ip_route(text)) for text in outputs.values())
    write_report(rows, out_path)
    for host, why in sorted(failures.items()):
        log.error('FAILED %s: %s', host, why)
    return len(outputs), len(failures)
Concern at 1000Answer
How many threads?Not 1000. A bounded pool of 20 to 50. Each worker holds an SSH session; the limits are the AAA/TACACS server's login rate, the control-plane CPU on small devices, and your own file descriptors. max_workers is the knob; start low and raise it while watching failure rate.
Rate limitingThe pool size already caps concurrency. For logins per second, sleep between submit calls or use a token bucket; for per-site limits, a threading.Semaphore per site acquired inside the worker.
One device hangsThe timeout inside run_command raises, with_retries tries again with backoff, then raises DeviceError, which lands in failures. The other 999 continue. The batch-level as_completed(timeout=) is the last line of defence.
Memory1000 route tables of a few hundred KB each fits, but do not rely on it. Parse in the worker and return the summary dict instead of the raw text, or write each output to Path(out) / f'{host}.txt' as it completes.
Partial resultsAlways produce the report for the hosts that worked plus a failure list with reasons. Exit code 3 for partial. Re-run only the failures: --hosts-from failures.txt. That re-run must be idempotent.
InventoryHosts come from a file or a source of truth (NetBox), never typed into the script. Validate the list before the first login.
Beyond a few thousandA framework that already does inventory, pools, retries and per-host results: nornir with netmiko or napalm, or an async stack (scrapli, asyncssh). Or split the fleet by site and run from a worker near each site.

The GIL in two sentences. CPython lets one thread execute Python bytecode at a time, so threads do not speed up pure Python computation. Waiting on a socket releases the GIL, so for SSH sessions threads give nearly linear speedup; if the bottleneck moved to CPU-heavy parsing you would use ProcessPoolExecutor for that stage.

asyncio in one paragraph. The same shape with one thread and an event loop: async def fetch(host) using an async SSH library (asyncssh, scrapli-asyncssh), asyncio.Semaphore(50) to bound concurrency, asyncio.wait_for(coro, timeout=15) per call, asyncio.gather(*tasks, return_exceptions=True) to collect results and exceptions together. It scales to more concurrent sessions than threads with less memory, but every library in the path must be async. In an interview, write the thread pool; mention asyncio as the next step.

Data structure design problems

The four design problems linked below are "automation mindset" in miniature: pick the containers that make each operation O(1) or O(log n), and state the guarantees out loud.

import random
from bisect import bisect_right

class RandomizedSet:                      # LC 380: list for O(1) random, dict for O(1) index lookup
    def __init__(self):
        self.vals, self.pos = [], {}
    def insert(self, v):
        if v in self.pos:
            return False
        self.pos[v] = len(self.vals); self.vals.append(v)
        return True
    def remove(self, v):
        if v not in self.pos:
            return False
        i, last = self.pos[v], self.vals[-1]
        self.vals[i], self.pos[last] = last, i       # move the last element into the hole
        self.vals.pop(); del self.pos[v]             # then pop: O(1), no shifting
        return True
    def getRandom(self):
        return random.choice(self.vals)

class TimeMap:                            # LC 981: timestamps per key arrive increasing, so bisect
    def __init__(self):
        self.ts, self.vals = {}, {}                  # key -> [timestamps], key -> [values]
    def set(self, key, value, t):
        self.ts.setdefault(key, []).append(t)
        self.vals.setdefault(key, []).append(value)
    def get(self, key, t):
        if key not in self.ts:
            return ''
        i = bisect_right(self.ts[key], t)            # number of timestamps <= t
        return self.vals[key][i - 1] if i else ''    # O(log k) for k versions of the key
ProblemContainersGuarantee to state
Insert Delete GetRandom O(1)list + dict of value to index; delete by swap with last and popall three O(1); random.choice needs a list, a dict cannot be indexed
Time Based Key-Value Storeper key two parallel lists (or a list of tuples); bisect_right on timestampsset O(1) amortised, get O(log k); relies on timestamps increasing, say so. Versioned config lookup: "what was this value at time t"
Design Authentication Managerdict token to expiry time; renew checks expiry first; countUnexpiredTokens scans or purges lazilygenerate/renew O(1); count O(n) unless you keep an OrderedDict in expiry order and pop from the front. Tokens with TTL is a session cache
Design Underground Systemdict id to (station, time) for open journeys; dict (start, end) to [total, count]all O(1); pop the check-in on check-out so memory is bounded by open journeys. Pairing start and end events is the flow-log pattern

Interview questions

1. json.dump raises TypeError: Object of type datetime is not JSON serializable. Fix?Pass default=str (or a function that formats the types you care about) so unencodable objects are converted instead of raising. Say that JSON has no date type, so the reader must parse the string back.
2. How do you make two JSON configs comparable?Load both, dumps each with sort_keys=True, indent=2, split into lines and feed difflib.unified_diff. Sorting keys removes ordering noise; indent puts one key per line so the diff is readable. Follow-up: lists are order-sensitive, so sort list members too when order does not matter.
3. Why newline='' when opening a file for the csv module?The csv module writes its own line endings and handles newlines inside quoted fields. If the file object also translates newlines you get \r\r\n on Windows (blank lines) and corrupted quoted fields. Same flag for reading.
4. How do you structure a script so it can be tested?Pure parse and decide functions that take data and return data, a thin main(argv) that does I/O, parse_args(argv=None) so tests can pass a list, and sys.exit(main()) under if __name__ == '__main__'. Then a self-test or pytest file calls the pure functions with a small sample and asserts.
5. Write a script that logs into 100 devices, runs sh ip route and parses hostnames and best routes. Now 1000.Serial loop first: each login about two seconds, so minutes for 100 and too long for 1000. The work is network waiting, so a ThreadPoolExecutor with 20 to 50 workers, each call with a connect and read timeout, bounded retries with backoff, failures collected per host without stopping the batch, results written as they complete, a CSV report plus a failure list, exit code for partial. Hostname from the prompt line, routes with one compiled regex per line. For more than a few thousand, nornir or an async stack, and split by site.
6. Python has the GIL. Why do threads help here at all?The GIL serialises Python bytecode, but blocking I/O releases it. An SSH session spends almost all its time waiting on the socket, so 20 threads give close to 20x. If parsing became CPU-bound I would move that stage to a process pool.
7. What happens if one device hangs and you have no timeout?That worker never returns. The pool keeps running the others, but as_completed never finishes and the with block waits forever at shutdown, so the report is never written. Threads cannot be killed, so the timeout has to be inside the call. I also put a batch-level timeout on as_completed as a backstop.
8. How do you retry safely?Only retry errors that can succeed later: timeouts, connection resets, a 503. Never retry authentication failures or parse errors. Bounded attempts, exponential backoff with a cap, random jitter so a thousand clients do not retry in lockstep, log every attempt, and make the operation idempotent so a retry after a half-finished attempt is harmless.
9. What does idempotent mean for a network change script?Running it twice leaves the device in the same state as running it once. "Ensure the VLAN exists" rather than "add the VLAN". It is what makes retries and re-runs after a crash safe, and it is why the script should read current state, compute a diff, and apply only the diff.
10. Why is a bare except: wrong, and what do you write instead?It catches everything including KeyboardInterrupt and SystemExit, so the script cannot be stopped and bugs are hidden. Catch the specific exceptions the operation can raise, as close to the call as possible. At a worker boundary except Exception is acceptable if you log it with log.exception and record the failure. Translate library errors with raise DeviceError(...) from e.
11. When does a generator pipeline not help?When a stage needs the whole input: sorting, groupby on unsorted data, computing a median, or anything that has to see the last line before emitting the first. Then memory is O(n) at that stage no matter how lazy the rest is, and I would say so and consider external sort or partitioning.
12. Insert, delete and getRandom in O(1). How?A list holds the values so random.choice is O(1), and a dict maps value to its index. Delete swaps the victim with the last element, updates that element's index, and pops the tail, so no shifting. Follow-up with duplicates allowed: dict of value to a set of indices.

Traps

Do before marking this topic done

  1. Labs 7 to 10 below, timed, blank editor: inventory join with UNKNOWN region, most-specific prefix match with ipaddress, merge two sorted logs as a stream, burst alerts with a per-host deque. Each one as a pure function plus a main that reads the fixtures and prints JSON.
  2. The 100-devices script end to end with the fake run_command: argparse (--hosts, --workers, --dry-run, -o), collect with a bounded pool and timeouts, with_retries, parse_show_ip_route, summarize, CSV report, failure list on stderr, exit code 3 on partial. Then say the 1000-device answer out loud in under a minute.
  3. Write parse_show_ip_route from memory against the sample above and make the ten-line self-test pass. Add a line your regex does not handle and decide: extend the regex or count it as unparsed?
  4. JSON config diff tool: two files in, sort_keys plus indent, difflib.unified_diff out, exit 0 when identical and 1 when different. Then make it accept JSON Lines and default=str a datetime field.
  5. Insert Delete GetRandom O(1) and Time Based Key-Value Store without looking. For each, say every operation's complexity before typing.

Practice: linked problems and labs

Logged attempts feed the tracker. Open the problem in a new tab, solve in a blank editor, then log honestly.

← 14 · BGP in depth · all topics · 16 · Meta's network: Clos fabrics, ECMP, BGP in the DC, FBOSS, backbone and MPLS →