Lesson 21 / 25
Monitoring What Matters
Track consumer lag, under-replicated partitions, request latency and disk usage.
Four signals
Alert on these first: (1) consumer lag per group (rising lag means consumers cannot keep up; also track time lag); (2) under-replicated partitions and offline partitions (non-zero means data is at risk or unavailable); (3) request latency and error rates on producers and brokers, and ISR shrink/expand events; (4) disk usage and network on brokers, since a full disk stops a broker. Export JMX metrics to Prometheus/Grafana or use a managed monitoring service, and watch lag trends, not only instantaneous values.
Watch, secure, scale
A production cluster needs monitoring, access control and a plan for capacity and upgrades.
Checking groups and under-replicated partitions
Both commands ship with Kafka. An empty result from --under-replicated-partitions is healthy. (Illustrative for a real cluster; the lag output format was shown earlier.)
kafka-consumer-groups.sh --bootstrap-server broker1:9092 --describe --group billing
kafka-topics.sh --bootstrap-server broker1:9092 --describe --under-replicated-partitions
kafka-topics.sh --bootstrap-server broker1:9092 --describe --unavailable-partitionsAlert on lag growth, not just size
A lag of 10,000 may be fine for a batch consumer and a crisis for a payments consumer. Alert when lag keeps growing for several minutes, with thresholds per consumer.
Quick check: What does a non-zero under-replicated partition count indicate?
- The cluster is idle
- Some replicas are not keeping up, so durability is reduced
- Consumers are too fast
- Compression is off
Answer
Some replicas are not keeping up, so durability is reduced — It means fewer in-sync copies exist than the replication factor, which needs attention.