Lesson 19 / 25

Kafka Connect

Move data between Kafka and databases, search engines and storage without writing consumer code.

Configuration instead of code

Kafka Connect is a framework for scalable, fault-tolerant data integration. Source connectors pull data into Kafka (for example a database's change log via Debezium change data capture), and sink connectors push topics out to systems such as Elasticsearch, S3, a warehouse or another database. You configure a connector with JSON and run it on a Connect cluster; it handles offsets, retries, scaling and schema conversion. Prefer a maintained connector to hand-written glue for common systems.

Move data and process it

Connect moves data in and out, and stream processors transform it as it flows.

Three tools: Connect, Streams, SQL.
Figure 6.1 — Connect, Streams and SQL.

A sink connector config (illustrative)

POST this JSON to the Connect REST API. Class names and options depend on the connector you install.

{
  "name": "orders-to-s3",
  "config": {
    "connector.class": "io.confluent.connect.s3.S3SinkConnector",
    "topics": "sales.order.placed.v1",
    "s3.bucket.name": "lake-raw",
    "storage.class": "io.confluent.connect.s3.storage.S3Storage",
    "format.class": "io.confluent.connect.s3.format.parquet.ParquetFormat",
    "flush.size": "10000",
    "tasks.max": "4"
  }
}

Quick check: What is Kafka Connect for?

  • Formatting code
  • Encrypting disks
  • Replacing brokers
  • Moving data between Kafka and other systems using connectors
Answer

Moving data between Kafka and other systems using connectors — Connect provides ready-made, scalable source and sink integrations.