All concepts

Kafka & the Commit Log

Not a queue that hands out messages and forgets them — an append-only log that keeps them, and readers who remember their own place in it.

Streaming & CDC · Intermediate · ~6 min

In plain English

A ship's logbook rather than an in-tray. Nothing is removed by being read; each reader keeps a bookmark and can turn back the pages.

Why it's worth your time

The 'log, not queue' inversion is what makes replay, fan-out and recovery possible at all — and it's the most common interview topic in streaming.

If you remember three things

  • Consumers own their offsets, so reading destroys nothing
  • Ordering exists per partition only, and the key picks the partition
  • Parallelism in a group is capped by the partition count

Overview

Kafka's central idea is that the broker stores an ordered, immutable log per partition and does not track who has read what. Consumers hold their own offset, so reading is a seek rather than a dequeue, and a message is not destroyed by being consumed. That inversion is why the same topic can feed a real-time service, a warehouse loader and a brand-new consumer replaying from the beginning, all at once and at their own speeds. A topic is split into partitions for parallelism, ordering is guaranteed only within a partition, and the key you choose decides which partition a record lands in — which makes key selection the most consequential design decision in the whole system.

In an interview

Kafka is a distributed append-only log. A topic is split into partitions; each partition is an ordered, immutable sequence, and consumers track their own offset rather than the broker tracking delivery. So messages aren't destroyed by reading, many independent consumer groups can read the same topic, and a new consumer can replay history. Ordering holds within a partition only, and the record key picks the partition.

Production defaults

Durability
acks=all, min.insync.replicas=2, idempotent producer
Key choice
the entity you need ordered, with enough distinct values to spread
Alerting
consumer lag in seconds behind, not in message count

What breaks

  • One consumer is always behind — A hot partition from a coarse key, or fewer partitions than consumers. Check per-partition lag.
  • Acknowledged writes lost after a broker failure — acks=1. Set acks=all with min.insync.replicas=2.

Watch it explained

Apache Kafka Fundamentals You Should Know — ByteByteGo, 4:54

Related