Lesson 25 / 25
Revision: Cheat Sheet and Self-Check
Review Kafka's concepts, settings and operational habits from the whole course.
Cheat sheet
Model: topic → partitions (ordered logs) → offsets; records have key/value/timestamp/headers; ordering only within a partition. Producers: key → murmur2 % partitions; acks=all, idempotence, linger.ms, compression; schemas with a registry. Consumers: group = one consumer per partition; commit after processing (at-least-once) and make handlers idempotent; lag = log-end − committed; rebalances; DLT for poison records. Topics: size partitions from throughput (can only increase; changes key mapping); retention vs compaction (tombstones); naming and ownership. Durability: rf=3, min.insync.replicas=2, acks=all, no unclean election; exactly-once with transactions inside Kafka. Ecosystem: Connect, Streams/Flink/ksqlDB. Ops: lag, URP, disk, latency; TLS + SASL + ACLs; rolling upgrades; capacity = rate × retention × rf.
Questions interviewers ask
Be ready to explain: how ordering works in Kafka and why it is per partition, how a consumer group scales, what at-least-once means and how you make processing idempotent, how acks, replication factor and min.insync.replicas combine, what happens when you add partitions, and how you would debug growing consumer lag.
Quick check: Consumer lag keeps growing. Which is a likely first check?
- Whether topics have long names
- Whether the font is readable
- Whether consumers are slow or stuck, or whether partitions/consumers are too few
- Deleting the topic
Answer
Whether consumers are slow or stuck, or whether partitions/consumers are too few — Lag grows when processing is slower than production, so check consumer speed, errors, rebalances and parallelism.
Quick check: Which setup survives one broker failure without losing acknowledged writes?
- rf=3, min.insync.replicas=2, acks=all
- rf=1, acks=1
- rf=2, min.insync.replicas=2, acks=all
- acks=0 with rf=3
Answer
rf=3, min.insync.replicas=2, acks=all — Three replicas with two required in sync let writes continue after one failure and keep data on two brokers.
Quick check: You want the processing of a record to be safe if it is delivered twice. What do you do?
- Add more partitions
- Commit before processing
- Disable retries on the broker
- Make the handler idempotent, for example upsert by a unique key
Answer
Make the handler idempotent, for example upsert by a unique key — Idempotent handling makes duplicate delivery harmless.