Every client setting that matters for performance trades latency against throughput by changing how much work one request carries. Producers and Delivery measured the producer side; the consumer has the mirror-image settings:
| Setting (Java name) | Default | For throughput | For latency |
|---|---|---|---|
| linger.ms, batch.size | 5 ms, 16 KB | 20-100 ms, 128 KB-1 MB | 0-5 ms |
| fetch.min.bytes | 1 byte | 64 KB-1 MB | 1 byte |
| fetch.max.wait.ms | 500 ms | 500 ms | 100-500 ms |
| max.poll.records | 500 | 1,000-5,000 | 100-500 |
A broker answers a fetch as soon as fetch.min.bytes are available or fetch.max.wait.ms has passed. With the default of one byte, a consumer of a quiet topic gets each event almost at once, at the price of one request per event. l0615_fetch.py sends 20 small order events a second for 10 seconds and reads them with the defaults and with a throughput setting (librdkafka calls the wait fetch.wait.max.ms):
trial("defaults", {})
trial("throughput", {"fetch.min.bytes": 65536, "fetch.wait.max.ms": 500})defaults p50 2 ms p99 34 ms 202 fetch requests for 200 events throughput p50 256 ms p99 504 ms 22 fetch requests for 200 events
Asking for 64 KB at a time cut fetch requests ninefold and raised the median delay from 2 ms to a quarter of a second, because each fetch now waited out its 500 ms for bytes that never came. For BookNest's analytics consumer, which feeds hourly reports, that is free capacity; for the packing service, where a customer waits, it is not.