JMX Metrics That Matter

Brokers and Java clients publish thousands of JMX MBeans; a handful decide whether the cluster is healthy:

The broker metrics worth an alert or a dashboard
MBean (abbreviated) Healthy value Meaning
ReplicaManager UnderReplicatedPartitions 0 Followers out of the ISR
ReplicaManager UnderMinIsrPartitionCount 0 acks=all writes failing
KafkaController OfflinePartitionsCount 0 Partitions with no leader
KafkaController ActiveControllerCount 1 in the cluster KRaft has one active controller
RequestHandlerAvgIdlePercent, NetworkProcessorAvgIdlePercent Above 0.3 Broker threads saturated
RequestMetrics TotalTimeMs (Produce, FetchConsumer) Stable p99 Request latency
BrokerTopicMetrics MessagesInPerSec, BytesInPerSec Expected trend Load per broker and topic

Kafka 129 ships kafka-jmx.sh to read MBeans over remote JMX, which the broker opens when KAFKA_JMX_PORT is set (9999 here). Read right after the shop producer loaded the 390,737 events:

monitoring/jmx_query.shShell
# jmx_query.sh: read broker MBeans once over remote JMX (with KAFKA_OPTS cleared)
docker exec -e KAFKA_OPTS= l2-kafka /opt/kafka/bin/kafka-jmx.sh --one-time true \
  --jmx-url service:jmx:rmi:///jndi/rmi://l2-kafka:9999/jmxrmi --report-format properties \
  --object-name kafka.server:type=ReplicaManager,name=UnderReplicatedPartitions \
  --object-name kafka.server:type=BrokerTopicMetrics,name=MessagesInPerSec \
  --object-name kafka.network:type=RequestMetrics,name=TotalTimeMs,request=Produce \
  | grep -E ':(Value|Count|OneMinuteRate|99thPercentile)='
Output
Trying to connect to JMX url: service:jmx:rmi:///jndi/rmi://l2-kafka:9999/jmxrmi
kafka.network:type=RequestMetrics,name=TotalTimeMs,request=Produce:99thPercentile=265.879999999
  99994
kafka.network:type=RequestMetrics,name=TotalTimeMs,request=Produce:Count=116
kafka.server:type=BrokerTopicMetrics,name=MessagesInPerSec:Count=390737
kafka.server:type=BrokerTopicMetrics,name=MessagesInPerSec:OneMinuteRate=5289.089632722467
kafka.server:type=ReplicaManager,name=UnderReplicatedPartitions:Value=0

The request metrics are histograms: 116 produce requests carried the whole load, with a 99th-percentile time of about 266 ms on this shared host. Meters hold a count plus 1-, 5- and 15-minute rates. Clients have their own MBeans, such as the producer's record-error-rate (Interceptors and Metrics) and the consumer's records-lag-max; librdkafka clients report the same through the statistics callback instead.