Lesson 17 / 25
Failure Scenarios and Leader Election
Walk through what happens when a broker, a network link or a consumer fails.
What to expect
Broker down: the controller elects a new leader for each affected partition from its ISR; clients refresh metadata and retry, so you see a short pause, not data loss (with the durable settings above). Follower falls behind: it is dropped from the ISR until it catches up, and acks=all writes then wait for fewer replicas, which is why min.insync.replicas matters. Network partition: clients that cannot reach the leader time out; leaders cut off from the controller are replaced. Consumer dies: its partitions are reassigned after the session times out, and a new owner resumes from the last committed offset (some records may be processed again). Producer retries: idempotence prevents duplicates. Practise these failures in a test cluster, not for the first time in an incident.
A failure drill checklist
Run each drill in a test environment and write down what you observed.
Drill Expect
stop one broker leader election within seconds; no data loss; producers retry
stop two brokers (rf=3,isr=2) writes to affected partitions fail loudly; reads may continue
kill a consumer rebalance; lag rises then recovers; duplicates possible
throttle a follower shrinks ISR; under-replicated partitions alert fires
restart the whole cluster rolling restart: one broker at a time, wait for ISR to recoverRolling restarts: one at a time
Restart brokers one at a time and wait until under-replicated partitions return to zero before touching the next, or you can drop below min.insync.replicas.
Quick check: A broker fails. Where does the new partition leader come from?
- From the producer
- From a random consumer
- From the in-sync replicas of that partition
- A new partition is created
Answer
From the in-sync replicas of that partition — Only in-sync replicas are eligible, so the new leader has all acknowledged records.