Retries and Errors

Retries, Timeouts and Error Handling

Producers retry retriable errors (a moving leader, too few in-sync replicas) with growing backoff until delivery.timeout.ms expires (120 s in Java, 300 s in librdkafka); non-retriable ones, such as an oversized record, fail at once. Here two of three brokers are stopped.

l0658_errors.py: one error retried until the deadline, one rejected at oncePython
import time
from confluent_kafka import Producer
p = Producer({"bootstrap.servers": "localhost:33194", "acks": "all",
              "delivery.timeout.ms": 15000, "retry.backoff.ms": 500})
start = time.perf_counter()
report = lambda err, msg: print(f"{msg.key().decode()}: {err.name() if err else 'ok'} "
                                f"after {time.perf_counter() - start:.1f} s")
p.produce("booknest.retry", '{"order_id": 7}', "7", on_delivery=report)  # 2 of 3 brokers down
try:
    p.produce("booknest.retry", "x" * 2_000_000, "big", on_delivery=report)  # over 1 MB
except Exception as e:                                # rejected at once, never retried
    print("big:", e.args[0].name(), "raised by produce()")
p.flush(30)
Output
big: MSG_SIZE_TOO_LARGE raised by produce()
7: _MSG_TIMED_OUT after 15.1 s

librdkafka rejected the 2 MB record itself (message.max.bytes). Order 7 was retried 18 times ("Not enough in-sync replicas" in the debug log), then timed out. Handle it in the callback: alert, park in an outbox, or stop.