What are common causes of consumer lag in Kafka and how do you address it?
Discover the common causes of Kafka consumer lag, how to measure it, and proven fixes like scaling consumers, adding partitions and tuning poll settings.
Expected Interview Answer
Consumer lag is the gap between the latest offset produced to a partition (log-end offset) and the offset a consumer group has committed; it grows when consumers process messages slower than producers write them, and you address it by scaling consumers, speeding up processing, or tuning fetch settings.
Common causes include too few consumers relative to partitions, slow per-message processing (heavy computation or slow downstream calls), uneven partitioning that hot-spots one consumer, long or blocking poll loops that trigger rebalances, and producer bursts outpacing consumers. You fix it by adding consumers up to the partition count, increasing partitions to raise parallelism, batching or parallelizing downstream work, tuning max.poll.records and max.poll.interval.ms, and monitoring lag with tools like kafka-consumer-groups or Burrow so you can react before it snowballs.
- Lower latency between event production and consumption
- Prevents unbounded backlog that risks data expiring under retention
- Better resource utilization by matching consumer parallelism to load
- Early detection through lag monitoring avoids outages
- More predictable end-to-end processing times
AI Mentor Explanation
Consumer lag is the pile of un-scored deliveries stacking up when the scorer writes slower than the bowler bowls. If overs pour in faster than one scorer can record, the backlog grows. You clear it by adding scorers to share the workload, simplifying what each records per ball, or splitting the book into sections, so the recorded score keeps pace with play on the field.
Step-by-Step Explanation
Step 1
Measure the lag
Use kafka-consumer-groups --describe or a tool like Burrow to see LAG per partition.
Step 2
Find the bottleneck
Check whether it is too few consumers, slow processing, or a hot partition.
Step 3
Scale consumers
Add consumer instances up to the partition count so each partition is handled by one consumer.
Step 4
Increase parallelism
Raise partition count and/or parallelize downstream work to process faster.
Step 5
Tune poll settings
Adjust max.poll.records and max.poll.interval.ms to avoid rebalances from long processing.
Step 6
Monitor continuously
Alert on rising lag so you react before the backlog risks hitting retention limits.
What Interviewer Expects
- A precise definition of lag as log-end offset minus committed offset
- That consumer parallelism is capped by the number of partitions
- Named causes: slow processing, too few consumers, hot partitions, rebalances
- Concrete fixes: scale out, add partitions, tune max.poll settings, optimize downstream
- Awareness of monitoring tools and retention-expiry risk
Common Mistakes
- Adding more consumers than partitions and expecting further speedup
- Ignoring that heavy per-message work or slow downstream calls is the real bottleneck
- Confusing lag with the total number of messages in the topic
- Letting long processing exceed max.poll.interval.ms and triggering repeated rebalances
Best Answer (HR Friendly)
“Consumer lag means the readers are falling behind the writers, so unread messages pile up. It usually happens because there are too few readers or each message takes too long to handle, and you fix it by adding readers or making the processing faster.”
Code Example
# See LAG per partition for a consumer group
kafka-consumer-groups.sh --bootstrap-server localhost:9092 \
--describe --group payments
# If lag is high and consumers already equal partitions,
# raise partition count to allow more parallel consumers
kafka-topics.sh --alter --topic payments \
--bootstrap-server localhost:9092 --partitions 12
# Then scale the consumer group up to 12 instancesFollow-up Questions
- Why can you not have more active consumers than partitions in a group?
- How do max.poll.records and max.poll.interval.ms influence lag and rebalances?
- What risks arise if lag grows beyond the topic's retention window?
- How would you detect and fix a single hot partition causing skewed lag?
- What is the difference between consumer lag and end-to-end latency?
MCQ Practice
1. Consumer lag is best defined as which quantity?
Lag is the difference between the newest offset produced and the offset the group has committed, per partition.
2. You already have as many consumers as partitions but lag keeps rising. What is a valid next step?
Parallelism is capped by partitions, so raising partitions (and consumers) increases throughput; extra consumers beyond partitions stay idle.
3. Which setting most directly helps avoid rebalances caused by slow processing?
max.poll.interval.ms bounds how long processing may take before the consumer is considered dead and a rebalance is triggered.
Flash Cards
What is consumer lag? — The gap between a partition's log-end offset and the consumer group's committed offset.
Main cap on consumer parallelism? — The number of partitions: extra consumers beyond partitions stay idle.
Common lag fixes? — Scale consumers, add partitions, optimize processing, tune max.poll settings.
Why monitor lag? — Rising lag can push a backlog past retention, risking permanent data loss.