Producer and Consumer Tuning

Every client setting that matters for performance trades latency against throughput by changing how much work one request carries. Producers and Delivery measured the producer side; the consumer has the mirror-image settings:

Client settings and the direction to move them
Setting (Java name) Default For throughput For latency
linger.ms, batch.size 5 ms, 16 KB 20-100 ms, 128 KB-1 MB 0-5 ms
fetch.min.bytes 1 byte 64 KB-1 MB 1 byte
fetch.max.wait.ms 500 ms 500 ms 100-500 ms
max.poll.records 500 1,000-5,000 100-500

A broker answers a fetch as soon as fetch.min.bytes are available or fetch.max.wait.ms has passed. With the default of one byte, a consumer of a quiet topic gets each event almost at once, at the price of one request per event. l0615_fetch.py sends 20 small order events a second for 10 seconds and reads them with the defaults and with a throughput setting (librdkafka calls the wait fetch.wait.max.ms):

listings/l0615_fetch.py (excerpt): the two consumers comparedPython
trial("defaults", {})
trial("throughput", {"fetch.min.bytes": 65536, "fetch.wait.max.ms": 500})
Output
defaults    p50    2 ms  p99   34 ms  202 fetch requests for 200 events
throughput  p50  256 ms  p99  504 ms  22 fetch requests for 200 events

Asking for 64 KB at a time cut fetch requests ninefold and raised the median delay from 2 ms to a quarter of a second, because each fetch now waited out its 500 ms for bytes that never came. For BookNest's analytics consumer, which feeds hourly reports, that is free capacity; for the packing service, where a customer waits, it is not.