What is a Kafka topic and how does partitioning work?
Learn what a Kafka topic is and how partitioning works: ordered partition logs, key hashing, consumer-group parallelism, and replication for fault tolerance.
Expected Interview Answer
A Kafka topic is a named, append-only log that groups related events, and partitioning splits that topic into multiple ordered logs called partitions so the data can be spread across brokers and consumed in parallel.
Each partition is an independent, ordered sequence of records, and Kafka only guarantees ordering within a single partition, not across the whole topic. A record's partition is chosen by hashing its key (or round-robin when no key is set), so all events with the same key land in the same partition and stay ordered. Partitions are the unit of parallelism and scaling: adding partitions lets more consumers in a group process the topic simultaneously, and each partition is replicated across brokers for fault tolerance.
- Parallelism: more partitions allow more concurrent consumers
- Ordering guaranteed per partition via keyed records
- Horizontal scaling by distributing partitions across brokers
- Fault tolerance through per-partition replication
- Even load distribution when keys are well chosen
AI Mentor Explanation
A topic is like the full match commentary, and partitions are like splitting that commentary across several scorers, one per batter. Every ball for a given batter always goes to the same scorer, so that batter's story stays in perfect order, while the scorers work in parallel to keep up with a fast game. No single scorer sees the whole match in order, but each batter's thread is flawless.
Step-by-Step Explanation
Step 1
Define a topic
Create a named topic and choose a partition count based on expected throughput and consumer parallelism.
Step 2
Producer sends a record
The producer supplies an optional key with each record's value.
Step 3
Choose a partition
Kafka hashes the key to pick a partition, or round-robins when no key is given.
Step 4
Append to the partition log
The record is appended to that partition and assigned a monotonically increasing offset.
Step 5
Consume in parallel
Each partition is read by at most one consumer in a group, so more partitions allow more parallel consumers.
Step 6
Replicate
Each partition has a leader and follower replicas across brokers for fault tolerance.
What Interviewer Expects
- Definition of a topic as a named append-only log
- Partitions as independent ordered logs and the unit of parallelism
- Ordering guaranteed only within a partition, not across the topic
- How keys determine partition assignment via hashing
- Relationship between partition count and consumer-group parallelism
- Awareness of replication for fault tolerance
Common Mistakes
- Claiming Kafka guarantees total ordering across a whole topic
- Forgetting that keyless records are distributed round-robin
- Thinking more consumers than partitions increases parallelism
- Ignoring that partition count is hard to reduce later
- Confusing partitions with replicas
Best Answer (HR Friendly)
“A Kafka topic is a labeled stream of related events, and partitioning splits that stream into several ordered lanes. Spreading events across lanes lets many workers process them at once, while events that share a key stay in the same lane so their order is preserved.”
Code Example
# Create a topic named 'orders' with 3 partitions
kafka-topics.sh --create \
--topic orders \
--partitions 3 \
--replication-factor 2 \
--bootstrap-server localhost:9092
# Produce keyed records: same key -> same partition -> ordered
kafka-console-producer.sh \
--topic orders \
--property parse.key=true \
--property key.separator=: \
--bootstrap-server localhost:9092
# > customer-7:placed
# > customer-7:paid # goes to the same partition as 'placed'Follow-up Questions
- How does Kafka decide which partition a record goes to?
- Why can't you have more active consumers than partitions in a group?
- How does replication differ from partitioning?
- What happens to ordering if you increase the partition count later?
- How do you choose a good partition key?
MCQ Practice
1. Kafka guarantees message ordering at which level?
Ordering is guaranteed only within a partition, not across all partitions of a topic.
2. How is a record's partition chosen when a key is provided?
Kafka hashes the key so all records with the same key consistently land in the same partition.
3. If a topic has 3 partitions, how many consumers in one group can actively read in parallel?
Each partition is read by at most one consumer per group, so parallelism is capped at the partition count.
Flash Cards
What is a Kafka topic? — A named, append-only log that groups related events, split into one or more partitions.
What is a partition? — An independent, ordered log within a topic; it is Kafka's unit of parallelism and scaling.
Where is ordering guaranteed? — Only within a single partition, not across the whole topic.
How is a partition chosen? — By hashing the record key, or round-robin when no key is set, so same-key records stay together.
Partitions vs consumers? — Each partition is read by at most one consumer in a group, so partition count caps parallelism.