Lesson 11 / 25
Commits and Delivery Semantics
Choose between at-most-once, at-least-once and exactly-once processing.
When you commit decides what can go wrong
If you commit before processing, a crash can lose records (at-most-once). If you process, then commit, a crash can reprocess some records (at-least-once), which is the usual default and requires your processing to be idempotent or de-duplicated. Exactly-once results need extra machinery (transactions, or an idempotent sink using the record's key or offset). Turn off auto-commit for important work and commit after the side effect succeeds, either synchronously or in batches.
A manual-commit consumer loop (illustrative)
Commit only after the database write succeeds. If the process dies in between, the record is read again, so save_order must tolerate duplicates (for example INSERT ... ON CONFLICT DO NOTHING).
from confluent_kafka import Consumer
c = Consumer({"bootstrap.servers": "localhost:9092", "group.id": "billing",
"enable.auto.commit": False, "auto.offset.reset": "earliest"})
c.subscribe(["orders"])
try:
while True:
msg = c.poll(1.0)
if msg is None: continue
if msg.error(): raise RuntimeError(msg.error())
save_order(msg.key(), msg.value()) # idempotent write
c.commit(message=msg) # commit AFTER success
finally:
c.close()Design for duplicates
Even with careful commits, retries and rebalances can deliver a record twice. Use a natural unique key (such as an order ID) so repeating the same record has no extra effect.
Quick check: Which order gives at-least-once processing?
- Commit twice
- Commit the offset, then process
- Never commit
- Process the record, then commit the offset
Answer
Process the record, then commit the offset — A crash between processing and commit causes a re-read, never a loss.