Lesson 8 / 25

Serialization and Schemas

Choose a record format and manage schema evolution with a schema registry.

Bytes in, contract needed

Kafka stores bytes; producers and consumers agree on how to turn objects into bytes. Plain JSON is easy to read but verbose and has no enforced schema. Avro, Protobuf and JSON Schema are compact and carry a defined schema. A schema registry (such as Confluent Schema Registry or Apicurio) stores versions of each schema, assigns IDs that are embedded in records, and enforces compatibility rules (for example backward-compatible: new readers can still read old data) so teams can add fields without breaking consumers. Choose compatibility deliberately and evolve schemas by adding optional fields with defaults, not by renaming or removing fields.

A backward-compatible change

Version 2 adds an optional coupon field with a default, so consumers on version 2 can still read records written with version 1. (Avro schema, illustrative.)

{
  "type": "record",
  "name": "OrderPlaced",
  "fields": [
    {"name": "order_id", "type": "string"},
    {"name": "amount",   "type": "double"},
    {"name": "coupon",   "type": ["null", "string"], "default": null}
  ]
}

Never reuse a field name with a new meaning

Changing what an existing field means silently breaks consumers even if the types still match. Add a new field instead and retire the old one gradually.

Quick check: Which schema change is generally backward-compatible?

  • Changing a field type from string to int
  • Renaming a required field
  • Adding an optional field with a default
  • Removing a required field
Answer

Adding an optional field with a default — A defaulted optional field lets new readers handle old records.