Lesson 19 / 25
Kafka Connect
Move data between Kafka and databases, search engines and storage without writing consumer code.
Configuration instead of code
Kafka Connect is a framework for scalable, fault-tolerant data integration. Source connectors pull data into Kafka (for example a database's change log via Debezium change data capture), and sink connectors push topics out to systems such as Elasticsearch, S3, a warehouse or another database. You configure a connector with JSON and run it on a Connect cluster; it handles offsets, retries, scaling and schema conversion. Prefer a maintained connector to hand-written glue for common systems.
Move data and process it
Connect moves data in and out, and stream processors transform it as it flows.
A sink connector config (illustrative)
POST this JSON to the Connect REST API. Class names and options depend on the connector you install.
{
"name": "orders-to-s3",
"config": {
"connector.class": "io.confluent.connect.s3.S3SinkConnector",
"topics": "sales.order.placed.v1",
"s3.bucket.name": "lake-raw",
"storage.class": "io.confluent.connect.s3.storage.S3Storage",
"format.class": "io.confluent.connect.s3.format.parquet.ParquetFormat",
"flush.size": "10000",
"tasks.max": "4"
}
}Quick check: What is Kafka Connect for?
- Formatting code
- Encrypting disks
- Replacing brokers
- Moving data between Kafka and other systems using connectors
Answer
Moving data between Kafka and other systems using connectors — Connect provides ready-made, scalable source and sink integrations.