Lesson 7 / 25

Batching and Compression

Tune linger.ms, batch.size and compression for throughput.

Trade a little latency for a lot of throughput

Producers do not send each record separately. They collect records per partition into batches. linger.ms is how long to wait for a batch to fill (0 sends immediately; 5 to 20 ms often improves throughput much), and batch.size is the maximum batch size in bytes. compression.type (lz4, zstd, snappy, gzip) compresses whole batches, reducing network and disk usage; lz4 and zstd are popular for good speed and ratio. Compression happens on the producer and stays on disk, and consumers decompress. Measure with your own data, since results depend on record size and content.

A producer in Python (illustrative)

This uses the confluent-kafka package (pip install confluent-kafka). It was not run here; the config keys are the standard names. flush() waits for outstanding sends.

from confluent_kafka import Producer

p = Producer({
    "bootstrap.servers": "localhost:9092",
    "acks": "all",
    "enable.idempotence": True,
    "linger.ms": 10,
    "compression.type": "lz4",
})

def on_delivery(err, msg):
    if err: print("failed:", err)
    else:   print(f"ok {msg.topic()}[{msg.partition()}]@{msg.offset()}")

for i in range(5):
    p.produce("orders", key=f"user-{i}", value=f"order {i}", on_delivery=on_delivery)
p.flush()

Always check delivery results

produce() is asynchronous. If you ignore the delivery callback or never call flush() before exiting, records can be lost without any error in your code.

Quick check: What does a small `linger.ms` such as 10 usually do?

  • Lets batches fill a little, raising throughput at small latency cost
  • Disables retries
  • Encrypts data
  • Deletes old records
Answer

Lets batches fill a little, raising throughput at small latency cost — Waiting briefly lets more records join a batch, improving efficiency.