Lesson 11 / 25

Commits and Delivery Semantics

Choose between at-most-once, at-least-once and exactly-once processing.

When you commit decides what can go wrong

If you commit before processing, a crash can lose records (at-most-once). If you process, then commit, a crash can reprocess some records (at-least-once), which is the usual default and requires your processing to be idempotent or de-duplicated. Exactly-once results need extra machinery (transactions, or an idempotent sink using the record's key or offset). Turn off auto-commit for important work and commit after the side effect succeeds, either synchronously or in batches.

A manual-commit consumer loop (illustrative)

Commit only after the database write succeeds. If the process dies in between, the record is read again, so save_order must tolerate duplicates (for example INSERT ... ON CONFLICT DO NOTHING).

from confluent_kafka import Consumer

c = Consumer({"bootstrap.servers": "localhost:9092", "group.id": "billing",
              "enable.auto.commit": False, "auto.offset.reset": "earliest"})
c.subscribe(["orders"])
try:
    while True:
        msg = c.poll(1.0)
        if msg is None: continue
        if msg.error(): raise RuntimeError(msg.error())
        save_order(msg.key(), msg.value())      # idempotent write
        c.commit(message=msg)                   # commit AFTER success
finally:
    c.close()

Design for duplicates

Even with careful commits, retries and rebalances can deliver a record twice. Use a natural unique key (such as an order ID) so repeating the same record has no extra effect.

Quick check: Which order gives at-least-once processing?

  • Commit twice
  • Commit the offset, then process
  • Never commit
  • Process the record, then commit the offset
Answer

Process the record, then commit the offset — A crash between processing and commit causes a re-read, never a loss.