A rule's condition must hold for its for: period before the alert fires, which filters out blips such as a broker restart; Alertmanager (not run here) then routes alerts to Slack 353 or PagerDuty (compare Slack and PagerDuty Alerts). BookNest pages for an unavailable cluster and files a ticket for a slow consumer:
# alerts.yml: BookNest's Kafka alert rules
groups:
- name: booknest-kafka
rules:
- alert: KafkaBrokerDown
expr: up{job="kafka"} == 0
for: 1m
labels: {severity: page}
- alert: KafkaNoActiveController
expr: sum(kafka_controller_kafkacontroller_activecontrollercount) != 1
for: 1m
labels: {severity: page}
- alert: KafkaOfflinePartitions
expr: sum(kafka_controller_kafkacontroller_offlinepartitionscount) > 0
for: 1m
labels: {severity: page}
- alert: KafkaUnderMinIsr
expr: sum(kafka_cluster_partition_underminisr) > 0
for: 5m
labels: {severity: page}
- alert: BookNestConsumerLagging
expr: sum by (group_id, topic_name) (kminion_kafka_consumer_group_topic_lag) > 10000
for: 2m
labels: {severity: ticket}
annotations:
summary: "{{ $labels.group_id }} is {{ $value }} events behind"# alerts.sh: validate the rules, wait past the 2-minute "for", list what fires
docker exec l2-prometheus promtool check rules /etc/prometheus/alerts.yml
sleep 150
curl -s localhost:32090/api/v1/alerts |
jq -r '.data.alerts[] | "\(.state) \(.labels.alertname) \(.annotations.summary // "")"'Checking /etc/prometheus/alerts.yml SUCCESS: 5 rules found firing BookNestConsumerLagging booknest-fraud is 389737 events behind
After two minutes the lag rule fired for the stopped group. Keep paging alerts few: offline partitions, no active controller, partitions under min.insync.replicas (acks=all writes failing), a broker not answering. Under-replicated partitions are normal during a rolling restart, so alert on them only after a long for:. Test the rules by stopping a broker in a staging cluster such as A Three-Node Quorum's quorum.