Lesson 16 / 25

min.insync.replicas and Availability

Combine replication factor, min.insync.replicas and acks=all to trade availability against durability.

The classic 3/2/all recipe

With acks=all, a write succeeds only when at least min.insync.replicas replicas (including the leader) have it; if fewer are in sync the producer gets an error rather than silently weakening the guarantee. The usual production recipe is replication factor 3, min.insync.replicas=2, acks=all: you can lose one broker and still write, and acknowledged data exists on at least two brokers. Losing two brokers makes the partition unavailable for writes (reads continue if a replica is in sync). That is the deliberate trade-off: consistency over availability. Keep unclean.leader.election.enable=false (the default) so an out-of-date replica never becomes leader and silently drops acknowledged records.

Surviving broker failures

Replication, ISR and the right producer settings decide whether acknowledged data survives a failure.

Three layers: replicate, acknowledge, elect.
Figure 5.1 — Replicate, acknowledge and elect.

How many failures can writes survive, run

I ran this for replication factor 3 and min.insync.replicas=2: writes work with 0 or 1 failed brokers (True, True) and stop at 2 or 3 failures (False, False).

def can_write(rf, min_isr, failed):
    return rf - failed >= min_isr

print([can_write(3, 2, f) for f in range(4)])

Output:

[True, True, False, False]

Setting it on a topic (illustrative)

Create a durable topic in one command.

kafka-topics.sh --bootstrap-server broker1:9092 --create --topic payments \
  --partitions 12 --replication-factor 3 --config min.insync.replicas=2

Quick check: With replication factor 3 and min.insync.replicas=2, how many broker failures can writes survive?

  • Three
  • Two
  • One
  • None
Answer

One — Writes need two in-sync replicas, so only one of the three may be down.