Brokers and Java clients publish thousands of JMX MBeans; a handful decide whether the cluster is healthy:
| MBean (abbreviated) | Healthy value | Meaning |
|---|---|---|
| ReplicaManager UnderReplicatedPartitions | 0 | Followers out of the ISR |
| ReplicaManager UnderMinIsrPartitionCount | 0 | acks=all writes failing |
| KafkaController OfflinePartitionsCount | 0 | Partitions with no leader |
| KafkaController ActiveControllerCount | 1 in the cluster | KRaft has one active controller |
| RequestHandlerAvgIdlePercent, NetworkProcessorAvgIdlePercent | Above 0.3 | Broker threads saturated |
| RequestMetrics TotalTimeMs (Produce, FetchConsumer) | Stable p99 | Request latency |
| BrokerTopicMetrics MessagesInPerSec, BytesInPerSec | Expected trend | Load per broker and topic |
Kafka 129 ships kafka-jmx.sh to read MBeans over remote JMX, which the broker opens when KAFKA_JMX_PORT is set (9999 here). Read right after the shop producer loaded the 390,737 events:
# jmx_query.sh: read broker MBeans once over remote JMX (with KAFKA_OPTS cleared)
docker exec -e KAFKA_OPTS= l2-kafka /opt/kafka/bin/kafka-jmx.sh --one-time true \
--jmx-url service:jmx:rmi:///jndi/rmi://l2-kafka:9999/jmxrmi --report-format properties \
--object-name kafka.server:type=ReplicaManager,name=UnderReplicatedPartitions \
--object-name kafka.server:type=BrokerTopicMetrics,name=MessagesInPerSec \
--object-name kafka.network:type=RequestMetrics,name=TotalTimeMs,request=Produce \
| grep -E ':(Value|Count|OneMinuteRate|99thPercentile)='Output
Trying to connect to JMX url: service:jmx:rmi:///jndi/rmi://l2-kafka:9999/jmxrmi kafka.network:type=RequestMetrics,name=TotalTimeMs,request=Produce:99thPercentile=265.879999999 99994 kafka.network:type=RequestMetrics,name=TotalTimeMs,request=Produce:Count=116 kafka.server:type=BrokerTopicMetrics,name=MessagesInPerSec:Count=390737 kafka.server:type=BrokerTopicMetrics,name=MessagesInPerSec:OneMinuteRate=5289.089632722467 kafka.server:type=ReplicaManager,name=UnderReplicatedPartitions:Value=0
The request metrics are histograms: 116 produce requests carried the whole load, with a 99th-percentile time of about 266 ms on this shared host. Meters hold a count plus 1-, 5- and 15-minute rates. Clients have their own MBeans, such as the producer's record-error-rate (Interceptors and Metrics) and the consumer's records-lag-max; librdkafka clients report the same through the statistics callback instead.