Lag is the distance between a partition's log-end offset and the group's committed offset: how many records a group has yet to process. kafka-consumer-groups.sh shows it on demand (kafka-consumer-groups), but brokers do not export it as a metric, so an exporter computes it from the group offsets: KMinion (https://github.com/redpanda-data/kminion 708 ) (MIT, 2.3.6) here. Two groups read the order topic:
# lag.sh: one group reads everything, one stops after 1,000 events; lag per group
K="docker exec -e KAFKA_OPTS= l2-kafka /opt/kafka/bin/kafka-console-consumer.sh"
T="--bootstrap-server l2-kafka:9092 --topic booknest.order-events --from-beginning"
$K $T --group booknest-audit --timeout-ms 5000 > /dev/null 2>&1 # reads to the end
$K $T --group booknest-fraud --max-messages 1000 > /dev/null 2>&1 # then stops
sleep 15 # KMinion and Prometheus catch up
curl -s localhost:32090/api/v1/query --data-urlencode \
'query=sum by (group_id) (kminion_kafka_consumer_group_topic_lag)' |
jq -r '.data.result[] | "\(.metric.group_id) lag \(.value[1])"'booknest-audit lag 0 booknest-fraud lag 389737
Lag in records hides speed: 389,737 events is seconds of work for Kafka Clients's consumers but hours for one that calls a slow service per event, so watch the trend. Burrow (https://github.com/linkedin/Burrow 3,966 ) (Apache 2.0, 1.9.6) judges groups over a sliding window instead of a threshold; kafka_exporter (https://github.com/danielqsj/kafka_exporter 2,538 ) (Apache 2.0) is an alternative to KMinion.