Producers retry retriable errors (a moving leader, too few in-sync replicas) with growing backoff until delivery.timeout.ms expires (120 s in Java, 300 s in librdkafka); non-retriable ones, such as an oversized record, fail at once. Here two of three brokers are stopped.
import time
from confluent_kafka import Producer
p = Producer({"bootstrap.servers": "localhost:33194", "acks": "all",
"delivery.timeout.ms": 15000, "retry.backoff.ms": 500})
start = time.perf_counter()
report = lambda err, msg: print(f"{msg.key().decode()}: {err.name() if err else 'ok'} "
f"after {time.perf_counter() - start:.1f} s")
p.produce("booknest.retry", '{"order_id": 7}', "7", on_delivery=report) # 2 of 3 brokers down
try:
p.produce("booknest.retry", "x" * 2_000_000, "big", on_delivery=report) # over 1 MB
except Exception as e: # rejected at once, never retried
print("big:", e.args[0].name(), "raised by produce()")
p.flush(30)Output
big: MSG_SIZE_TOO_LARGE raised by produce() 7: _MSG_TIMED_OUT after 15.1 s
librdkafka rejected the 2 MB record itself (message.max.bytes). Order 7 was retried 18 times ("Not enough in-sync replicas" in the debug log), then timed out. Handle it in the callback: alert, park in an outbox, or stop.