BookNest's runbook moves more than records:
Inventory topics, groups, ACLs, schemas, connectors, client settings and What Managed Takes Off You's ties.
Prepare the target: principals and ACLs, a registry with the same schema IDs (Schema Registry's import mode, or an empty registry fed the _schemas topic), connectors created but stopped.
Mirror for days, watching MM2's replication latency.
Cut over in a quiet hour: stop producers, let consumers catch up, wait for equal end offsets, stop MM2 and the consumers, start producers on the target, then consumers.
Keep the old cluster read-only for a week as the way back.
migration/cutover.sh rehearses step 4: the shop producer (produce.py FIRST LAST, Producing Order Events's for a range of events) writes 80,000 more to l2-old while MM2 copies them, analytics catches up, and both move.
# cutover.sh: move BookNest from old to new while MirrorMaker 2 runs (after mirror.sh)
. ./lib.sh # k CLUSTER TOOL ARGS runs a Kafka tool; ends CLUSTER
PY=/home/dev/v7-l2/venv-kafka/bin/python; G="--group booknest-analytics"
lag() { k $1 consumer-groups --describe $G 2>/dev/null |
awk '$3 ~ /^[0-9]/ {s += $6} END {print s}'; }
read_all() { k $1 console-consumer --topic booknest.order-events $G --timeout-ms 10000 \
2>/dev/null | wc -l; }
# 1. Live traffic still goes to old while MM2 copies it; analytics catches up there
$PY produce.py localhost:32093 300001 380000
echo "analytics read on old: $(read_all old)"
# 2. Freeze: producers stopped; wait until new holds every record and the synced offsets
for i in $(seq 60); do [ "$(ends new)" = "$(ends old)" ] && break; sleep 5; done; sleep 15
for c in old new; do echo "$c: end offsets $(ends $c), analytics lag $(lag $c)"; done
# 3. Stop MM2, move the producer to new, and resume analytics there
docker stop l2-mm2 >/dev/null
$PY produce.py localhost:32094 380001 390737
echo "analytics read on new: $(read_all new)"localhost:32093: 80,000 sent, 0 failed, 0 undelivered analytics read on old: 180000 old: end offsets 126869 126858 126273, analytics lag 0 new: end offsets 126869 126858 126273, analytics lag 0 localhost:32094: 10,737 sent, 0 failed, 0 undelivered analytics read on new: 10737
MM2 translated the group's final position exactly (lag 0 on both clusters), and on l2-new it read precisely the 10,737 events the shop wrote there: nothing lost or read twice, and l2-new holds all 390,737. Exact is not guaranteed: in an earlier run the position trailed by 338 records, about one offset-sync interval per partition, and 338 events were read twice. Catching up first keeps the overlap that small; consumers that skip stored event_ids (Delivery Semantics) make it harmless.
MM2 writes translated offsets only while the group has no members on the target, so start it there only after the last sync. And keep the producers' switch a configuration change that reverses as easily, with an MM2 flow back ready for whatever the target received before a rollback.