Lesson 1 / 25

What Kafka Is

Describe Kafka as a distributed, replicated, append-only event log and say when to use it.

Events in, events out, events kept

Apache Kafka is a distributed platform for streaming events. Systems called producers write records to named topics; systems called consumers read them. Unlike a classic queue that deletes a message once it is consumed, Kafka stores records on disk for a configured time or size and lets many independent consumers read the same data, each at its own position. This makes it useful for decoupling services, moving data between systems (log shipping, change data capture), feeding analytics and building real-time pipelines. It is not a database for random lookups and not a task queue with per-message acknowledgements.

A durable log everyone can read

Producers append events to a topic; consumers read them at their own pace; the log keeps them for as long as you configure.

Four parts: producers, topic log, brokers, consumers.
Figure 1.1 — Producers, topic log, brokers and consumers.

A newspaper archive

A newspaper prints each edition once and keeps a bound archive. Readers pick up from any date, at their own speed, without removing the edition for others. Kafka's topic is that archive.

Use the right tool

For a simple background job queue (send one email, ack it), RabbitMQ or a cloud queue may be simpler. Choose Kafka when you need replay, many consumers, high throughput or event streaming.

Quick check: What happens to a record after one consumer reads it?

  • It moves to another topic automatically
  • It is deleted immediately
  • It stays in the topic until retention removes it, and others can still read it
  • It is encrypted
Answer

It stays in the topic until retention removes it, and others can still read it — Kafka is a log: reading does not remove data, so many consumers can replay it.