A partition is a series of segments. Only the newest, the active segment, takes writes; a new one starts when it reaches segment.bytes (1 GiB) or its records span more than segment.ms (7 days). Each segment is a .log of batches named after its first offset, an .index from offsets to byte positions and a .timeindex from timestamps to offsets. The indexes are sparse, one entry per index.interval.bytes (4 KiB) or so, and stay in memory. The listing loads the history with load_events.py, a short confluent-kafka producer (Producers and Delivery) that keys events by order_id and sets each record's timestamp to the event time.
# The full order history in a topic that keeps it forever, then partition 0's files on broker 4
K="docker exec l3-c1 /opt/kafka/bin"
$K/kafka-topics.sh --bootstrap-server l3-b4:9092 --create --topic booknest.order-events \
--config retention.ms=-1 # never delete by time
/home/dev/v7-l3/kafka-venv/bin/python listings/load_events.py booknest.order-events
P=/var/lib/kafka/data/booknest.order-events-0
docker exec l3-b4 ls $P | sed 's/.*\.//' | sort | uniq -c | sort -rn | xargs
files() { docker exec l3-b4 ls -l $P | awk 'NR > 1 {printf "%9s %s\n", $5, $9}'; }
files | head -3 # the first segment
files | tail -6 | head -2 # the active segment's index and log390737 events sent, 0 still queued
69 timeindex 69 log 69 index 68 snapshot 1 metadata 1 leader-epoch-checkpoint
16 00000000000000000000.index
341745 00000000000000000000.log
24 00000000000000000000.timeindex
10485760 00000000000000128596.index
42359 00000000000000128596.logPartition 0 holds about 12 MB, yet 69 segments: segment.ms is measured on record timestamps, so 18 months of history rolled a segment about every week of event time. The first segment's 342 KB need two index entries, one per batch. The active index is preallocated at segment.index.bytes (10 MiB) and trimmed on roll. To read offset N, the broker finds the segment by its base offset, binary-searches its index and scans a few kilobytes of log.