Alerting

Alerting on Broker and Cluster Health

A rule's condition must hold for its for: period before the alert fires, which filters out blips such as a broker restart; Alertmanager (not run here) then routes alerts to Slack 353 or PagerDuty (compare Slack and PagerDuty Alerts). BookNest pages for an unavailable cluster and files a ticket for a slow consumer:

monitoring/prometheus/alerts.ymlYAML
# alerts.yml: BookNest's Kafka alert rules
groups:
  - name: booknest-kafka
    rules:
      - alert: KafkaBrokerDown
        expr: up{job="kafka"} == 0
        for: 1m
        labels: {severity: page}
      - alert: KafkaNoActiveController
        expr: sum(kafka_controller_kafkacontroller_activecontrollercount) != 1
        for: 1m
        labels: {severity: page}
      - alert: KafkaOfflinePartitions
        expr: sum(kafka_controller_kafkacontroller_offlinepartitionscount) > 0
        for: 1m
        labels: {severity: page}
      - alert: KafkaUnderMinIsr
        expr: sum(kafka_cluster_partition_underminisr) > 0
        for: 5m
        labels: {severity: page}
      - alert: BookNestConsumerLagging
        expr: sum by (group_id, topic_name) (kminion_kafka_consumer_group_topic_lag) > 10000
        for: 2m
        labels: {severity: ticket}
        annotations:
          summary: "{{ $labels.group_id }} is {{ $value }} events behind"
monitoring/alerts.shShell
# alerts.sh: validate the rules, wait past the 2-minute "for", list what fires
docker exec l2-prometheus promtool check rules /etc/prometheus/alerts.yml
sleep 150
curl -s localhost:32090/api/v1/alerts |
  jq -r '.data.alerts[] | "\(.state)  \(.labels.alertname)  \(.annotations.summary // "")"'
Output
Checking /etc/prometheus/alerts.yml
  SUCCESS: 5 rules found
firing  BookNestConsumerLagging  booknest-fraud is 389737 events behind

After two minutes the lag rule fired for the stopped group. Keep paging alerts few: offline partitions, no active controller, partitions under min.insync.replicas (acks=all writes failing), a broker not answering. Under-replicated partitions are normal during a rolling restart, so alert on them only after a long for:. Test the rules by stopping a broker in a staging cluster such as A Three-Node Quorum's quorum.