# Failure Scenarios and Leader Election — Apache Kafka: Event Streaming from Basics to Production

Source: https://www.geekswithgeeks.com/en/kafka/rel-failure-scenarios

> Walk through what happens when a broker, a network link or a consumer fails.

## What to expect

**Broker down**: the controller elects a new leader for each affected partition from its ISR; clients refresh metadata and retry, so you see a short pause, not data loss (with the durable settings above). **Follower falls behind**: it is dropped from the ISR until it catches up, and `acks=all` writes then wait for fewer replicas, which is why `min.insync.replicas` matters. **Network partition**: clients that cannot reach the leader time out; leaders cut off from the controller are replaced. **Consumer dies**: its partitions are reassigned after the session times out, and a new owner resumes from the last committed offset (some records may be processed again). **Producer retries**: idempotence prevents duplicates. Practise these failures in a test cluster, not for the first time in an incident.

## A failure drill checklist

Run each drill in a test environment and write down what you observed.

```text
Drill                         Expect
stop one broker               leader election within seconds; no data loss; producers retry
stop two brokers (rf=3,isr=2) writes to affected partitions fail loudly; reads may continue
kill a consumer               rebalance; lag rises then recovers; duplicates possible
throttle a follower           shrinks ISR; under-replicated partitions alert fires
restart the whole cluster     rolling restart: one broker at a time, wait for ISR to recover
```

## Rolling restarts: one at a time

Restart brokers one at a time and wait until under-replicated partitions return to zero before touching the next, or you can drop below `min.insync.replicas`.

**Quiz:** A broker fails. Where does the new partition leader come from?

- [ ] From the producer
- [ ] From a random consumer
- [x] From the in-sync replicas of that partition
- [ ] A new partition is created

*Answer:* From the in-sync replicas of that partition. Only in-sync replicas are eligible, so the new leader has all acknowledged records.
