The in-sync replica set (ISR) holds the replicas caught up with the leader; a follower that has not reached the leader's log end within replica.lag.time.max.ms (30 seconds), or whose broker is fenced, drops out. An acks=all write is acknowledged once every ISR member has it, and min.insync.replicas is the floor below which such writes fail. This cluster uses three replicas and a minimum ISR of two. The listing stops partition 0's leader, then another broker, and writes after each step.
# Stop partition 0's leader, then one more broker; after each step, write with acks=all
K="docker exec -i l3-c1 /opt/kafka/bin" # tools run on a controller
B="--bootstrap-server l3-b4:9092,l3-b5:9092,l3-b6:9092"
T="--topic booknest.ops-check"
p0() { $K/kafka-topics.sh $B --describe $T 2>/dev/null | awk -F'\t' '$3 == "Partition: 0"'; }
state() { p0 | awk -F'\t' '{print $4, $5, $6, $7}'; }
leader() { p0 | awk -F'\t' '{print $4}' | tr -dc 0-9; }
send() { echo "7|$1" | $K/kafka-console-producer.sh $B $T --command-property acks=all \
--reader-property parse.key=true --reader-property key.separator='|' \
--command-property delivery.timeout.ms=8000 \
--command-property request.timeout.ms=5000 2>&1 |
grep -o 'NotEnoughReplicas[A-Za-z]*' | head -1 | grep . || echo "written"; }
$K/kafka-topics.sh $B --create $T >/dev/null 2>&1 # 3 partitions, RF 3, min ISR 2
echo "all up: $(state) -> $(send first)"
first=$(leader); docker stop l3-b$first >/dev/null; sleep 5
second=$(leader); third=$((4 + 5 + 6 - first - second))
echo "$first stopped: $(state) -> $(send second)"
docker stop l3-b$third >/dev/null; sleep 5
echo "$third stopped too: $(state) -> $(send third)"
docker start l3-b$first l3-b$third >/dev/null
until state | grep -q 'Isr: .,.,.'; do sleep 2; done # until both have caught up
echo "both back: $(state)"all up: Leader: 4 Replicas: 4,5,6 Isr: 4,5,6 Elr: -> written 4 stopped: Leader: 5 Replicas: 4,5,6 Isr: 5,6 Elr: -> written 6 stopped too: Leader: 5 Replicas: 4,5,6 Isr: 5 Elr: 6 -> NotEnoughReplicasException both back: Leader: 5 Replicas: 4,5,6 Isr: 4,5,6 Elr:
Broker 5 took over within seconds and the write landed on two copies. With one replica left the partition stayed readable, but Kafka 129 refused acks=all writes that one disk failure could lose. The returning brokers caught up and rejoined the ISR. Three replicas with a minimum ISR of two survive one broker failure without losing writes.