Lesson 17 / 25

Failure Scenarios and Leader Election

Walk through what happens when a broker, a network link or a consumer fails.

What to expect

Broker down: the controller elects a new leader for each affected partition from its ISR; clients refresh metadata and retry, so you see a short pause, not data loss (with the durable settings above). Follower falls behind: it is dropped from the ISR until it catches up, and acks=all writes then wait for fewer replicas, which is why min.insync.replicas matters. Network partition: clients that cannot reach the leader time out; leaders cut off from the controller are replaced. Consumer dies: its partitions are reassigned after the session times out, and a new owner resumes from the last committed offset (some records may be processed again). Producer retries: idempotence prevents duplicates. Practise these failures in a test cluster, not for the first time in an incident.

A failure drill checklist

Run each drill in a test environment and write down what you observed.

Drill                         Expect
stop one broker               leader election within seconds; no data loss; producers retry
stop two brokers (rf=3,isr=2) writes to affected partitions fail loudly; reads may continue
kill a consumer               rebalance; lag rises then recovers; duplicates possible
throttle a follower           shrinks ISR; under-replicated partitions alert fires
restart the whole cluster     rolling restart: one broker at a time, wait for ISR to recover

Rolling restarts: one at a time

Restart brokers one at a time and wait until under-replicated partitions return to zero before touching the next, or you can drop below min.insync.replicas.

Quick check: A broker fails. Where does the new partition leader come from?

  • From the producer
  • From a random consumer
  • From the in-sync replicas of that partition
  • A new partition is created
Answer

From the in-sync replicas of that partition — Only in-sync replicas are eligible, so the new leader has all acknowledged records.