Lesson 23 / 25
Capacity, Upgrades and Managed Services
Plan storage and throughput, roll out upgrades safely and decide between self-managed and managed Kafka.
Plan, then keep headroom
Estimate storage (data per day × retention × replication factor, plus ~20-30% headroom), network (producers in + replication + consumers out) and partitions per broker. Upgrade with rolling restarts, one broker at a time, following the release notes and upgrade order, and test in staging first. Decide who runs the cluster: self-managed (full control, requires expertise for upgrades, security and on-call) or a managed service (Confluent Cloud, Amazon MSK, Azure Event Hubs for Kafka, Redpanda Cloud and others) which trades some control and cost for less operational work. Spread brokers across availability zones and enable rack awareness so replicas live in different zones.
A capacity worksheet
Numbers are examples. The point is to calculate before buying hardware or choosing a plan.
Ingress 20 MB/s average, 60 MB/s peak
Replication factor 3 -> inter-broker traffic ~ 2 x ingress
Consumers 3 groups -> egress ~ 3 x ingress
Retention 7 days -> 20 MB/s x 86,400 s x 7 x 3 replicas ~ 36 TB raw, +25% headroom
Brokers 6 across 3 AZs (rack awareness on)
Partitions ~1,000-2,000 per broker is a comfortable range for most clustersQuick check: Why upgrade brokers one at a time?
- To keep enough in-sync replicas available so writes continue
- Brokers cannot restart together
- It makes upgrades free
- Only for managed services
Answer
To keep enough in-sync replicas available so writes continue — Restarting all brokers at once would drop below min.insync.replicas and halt writes.