High Watermark and Epochs

The High Watermark and the Leader Epoch

Each replica has its own log end offset (LEO). The lowest LEO among the ISR is the high watermark (HW): records below it are on every in-sync replica, so they are committed, and only they are visible to consumers.

Three replicas of one partition: the high watermark is the offset every in-sync replica has reached
Three replicas of one partition: the high watermark is the offset every in-sync replica has reached

The HW alone could not stop a returning replica from truncating too much, or a former leader from keeping records no one else has. KIP-101 added the leader epoch, raised at every leader change and stamped on each batch; each replica's leader-epoch-checkpoint maps epochs to first offsets, so a returning replica asks the leader where its last epoch ended and truncates exactly there.

Leader epochs and high watermarks of partition 0 on every broker
# Partition 0's leader epoch cache and high watermark on each broker, and its batches' epochs
D=/var/lib/kafka/data
for b in 4 5 6; do
  epochs=$(docker exec l3-b$b tail -n +3 $D/booknest.ops-check-0/leader-epoch-checkpoint)
  hw=$(docker exec l3-b$b grep 'booknest.ops-check 0 ' $D/replication-offset-checkpoint)
  echo "broker $b: epoch/start offset: $(echo $epochs | xargs)  high watermark: ${hw##* }"
done
docker exec l3-b4 /opt/kafka/bin/kafka-dump-log.sh \
  --files $D/booknest.ops-check-0/00000000000000000000.log |
  awk '/baseOffset/ {print "offset", $2, "written in leader epoch", $16}'
Output
broker 4: epoch/start offset: 0 0 1 1  high watermark: 2
broker 5: epoch/start offset: 0 0 1 1  high watermark: 2
broker 6: epoch/start offset: 0 0 1 1  high watermark: 2
offset 0 written in leader epoch 0
offset 1 written in leader epoch 1

The first write went into epoch 0 under broker 4, the second into epoch 1 under broker 5, and the refused third left nothing; all replicas agree.