Capacity Planning

Capacity Planning on a Single-Node Cluster

Capacity planning turns business numbers into disk, network and CPU. Disk is events a day x bytes per event on disk x retention in days x replication factor, plus headroom for a broker to fail over; network in is the produce rate times the replication factor, and network out the produce rate times the number of consumer groups plus the replicas' copies. For BookNest the sample data gives the inputs:

BookNest's capacity inputs
Input BookNest (sample data) Source
Events 390,737 over 546 days, about 716 a day order_events.jsonl
Peaks 820 in a day, 60 in an hour Same file
Size on disk 94 bytes of JSON; 8.1 MB per 18 months with zstd 126 Compression Codecs
One broker's ceiling 220,000-310,000 keyed events a second Sizing Partitions

At these rates a single broker is idle: the busiest hour asked for one event a minute, against a measured ceiling more than 10,000,000 times higher, and the whole 18-month history with zstd fits in 8 MB, 24 MB with three replicas. Even a hundredfold growth and every event kept forever would need under 2 GB of disk a year. For BookNest the limit is not throughput at all but availability: one node is one failure away from an outage, so production gets three brokers (or three controllers and three brokers, A Three-Node Quorum) because of fault tolerance, not load.

A single development node like l3-kafka needs a 512 MB-1 GB heap (Kafka with Other Services) and free RAM for the page cache. As traffic grows, watch disk use against retention, request-handler idle percentage (Throughput and Latency) and consumer lag (Monitoring and Kafka UIs), and test each change with the tools of Benchmarking first.