Lesson 21 / 25

Monitoring What Matters

Track consumer lag, under-replicated partitions, request latency and disk usage.

Four signals

Alert on these first: (1) consumer lag per group (rising lag means consumers cannot keep up; also track time lag); (2) under-replicated partitions and offline partitions (non-zero means data is at risk or unavailable); (3) request latency and error rates on producers and brokers, and ISR shrink/expand events; (4) disk usage and network on brokers, since a full disk stops a broker. Export JMX metrics to Prometheus/Grafana or use a managed monitoring service, and watch lag trends, not only instantaneous values.

Watch, secure, scale

A production cluster needs monitoring, access control and a plan for capacity and upgrades.

Three duties: monitor, secure, plan.
Figure 7.1 — Monitor, secure and plan.

Checking groups and under-replicated partitions

Both commands ship with Kafka. An empty result from --under-replicated-partitions is healthy. (Illustrative for a real cluster; the lag output format was shown earlier.)

kafka-consumer-groups.sh --bootstrap-server broker1:9092 --describe --group billing
kafka-topics.sh --bootstrap-server broker1:9092 --describe --under-replicated-partitions
kafka-topics.sh --bootstrap-server broker1:9092 --describe --unavailable-partitions

Alert on lag growth, not just size

A lag of 10,000 may be fine for a batch consumer and a crisis for a payments consumer. Alert when lag keeps growing for several minutes, with thresholds per consumer.

Quick check: What does a non-zero under-replicated partition count indicate?

  • The cluster is idle
  • Some replicas are not keeping up, so durability is reduced
  • Consumers are too fast
  • Compression is off
Answer

Some replicas are not keeping up, so durability is reduced — It means fewer in-sync copies exist than the replication factor, which needs attention.