Lesson 6 / 25

Acknowledgements, Retries and Idempotence

Configure acks, retries and the idempotent producer to avoid loss and duplicates.

How sure do you want to be?

The acks setting controls when a write counts as successful: 0 (do not wait; fastest, may lose data), 1 (wait for the leader only), all (wait for all in-sync replicas; strongest). Combined with the topic setting min.insync.replicas, acks=all protects acknowledged data from a broker failure. The producer retries transient errors, which can create duplicates if a retry follows a write that actually succeeded. Setting enable.idempotence=true gives each producer an ID and sequence numbers so the broker discards duplicates within a partition; in current Kafka clients this is on by default together with acks=all. Use delivery.timeout.ms to bound the total time spent retrying.

A durable producer config

Illustrative property names from the Java/librdkafka clients. max.in.flight.requests.per.connection up to 5 keeps ordering with idempotence.

acks=all
enable.idempotence=true
retries=2147483647
delivery.timeout.ms=120000
max.in.flight.requests.per.connection=5
linger.ms=10
compression.type=lz4

How the broker drops duplicates, run

I ran a small model of the sequence check: each producer ID must send sequence numbers 0, 1, 2…; a repeated number is ignored as a duplicate and a gap is an error.

seen = {}
def accept(pid, seq):
    last = seen.get(pid, -1)
    if seq == last + 1:
        seen[pid] = seq
        return "append"
    if seq <= last:
        return "duplicate-ignored"
    return "out-of-order-error"

print([accept(7, s) for s in (0, 1, 1, 2, 4)])

Output:

['append', 'append', 'duplicate-ignored', 'append', 'out-of-order-error']

Quick check: Which setting makes the producer wait for all in-sync replicas?

  • compression.type=none
  • acks=0
  • linger.ms=0
  • acks=all
Answer

acks=all — acks=all acknowledges only after every in-sync replica has the record.