Distributed systems

Kafka consumer group stuck rebalancing

Written and reviewed by Sahil Srivastav

Apache KafkaConsumer groupLiveness
[Consumer clientId=consumer-orders-4, groupId=orders] Member consumer-orders-4-9f1c2a sending LeaveGroup request to coordinator broker-2:9092 (id: 2147483645 rack: null) due to consumer poll timeout has expired. This means the time between subsequent calls to poll() was longer than the configured max.poll.interval.ms, which typically implies that the poll loop is spending too much time processing messages.
[Consumer clientId=consumer-orders-4, groupId=orders] (Re-)joining group
[Consumer clientId=consumer-orders-4, groupId=orders] Successfully joined group with generation Generation{generationId=4187, memberId=..., protocol=range}

$ kafka-consumer-groups.sh --bootstrap-server broker:9092 --describe --group orders --state
GROUP   COORDINATOR (ID)   ASSIGNMENT-STRATEGY   STATE                #MEMBERS
orders  broker-2:9092 (2)  range                 PreparingRebalance   11

What this error actually means

A rebalance is a stop-the-world event for a consumer group. With the eager protocol, every member gives up every partition, the coordinator computes a new assignment, and only then does processing resume. Nothing is consumed in between. A group that rebalances every few seconds is therefore not slow — it is stopped, and the lag graph will show it.

The loop is self-sustaining, which is what makes it feel unfixable. One member is late calling `poll()`, so it is evicted and the group rebalances. The rebalance stalls every other member for its duration, so they too fall behind their own deadlines. When processing restarts, each member now has a bigger backlog and returns fuller batches, which take longer, which breaches the interval again. Each iteration makes the next one more likely.

The generation id is the counter that proves it. Every completed rebalance increments `generationId`. If that number climbs by tens per minute in your logs while no deployment, scale event, or node failure is happening, the group is thrashing rather than reacting to change — and the cause is inside the consumers, not in the cluster.

Causes, most common first

  1. 1One member with a pathological batch keeps triggering the loop. Only one consumer needs to exceed `max.poll.interval.ms` to stall the whole group. That member is usually the one that owns a hot partition or a large-payload key. Averaged metrics hide it completely — the group-level average looks fine while a single instance is always late.
  2. 2Rolling restarts with the eager assignor and a long session timeout. Each pod that leaves and rejoins causes two rebalances. With ten pods, a default `session.timeout.ms`, and no static membership, a routine deploy produces twenty stop-the-world events, and if `group.initial.rebalance.delay.ms` is low the coordinator will not batch them.
  3. 3A liveness probe that restarts healthy-but-busy consumers. An HTTP readiness endpoint served by the same thread pool that the consumer starves will time out under load. The orchestrator kills the pod, the group rebalances, lag grows, the next pod is starved, and the cluster now restarts consumers as fast as they can join.
  4. 4Members with mismatched configuration or assignors. During a partial deploy, members advertising different `partition.assignment.strategy` lists cannot agree on a protocol; some clients fail to join and retry in a loop. The same happens when one deployment sets a much shorter `max.poll.interval.ms` than the rest.
  5. 5More consumers than partitions, repeatedly re-shuffled. Idle members still participate in every rebalance. Scaling a twelve-partition topic to thirty consumers adds eighteen members that do no work but make every assignment computation and every join round larger and slower.

When you see it

  • `generationId` in the client logs increases continuously with no deploys
  • Lag grows steadily even though every consumer process is up and using CPU
  • `--state` reports `PreparingRebalance` or `CompletingRebalance` on nearly every check
  • Adding more consumer instances makes the situation worse rather than better
  • `CommitFailedException` accompanies it, and downstream duplicates accumulate
  • The group recovers briefly after a full restart of all members, then degrades again within minutes

How to diagnose it

Step 1

Count generations per minute

This is the decisive measurement and it takes one command. A healthy group increments the generation only when membership genuinely changes. Continuous increments mean a loop, and the member id on the `LeaveGroup` line names the instance to investigate first.

grep -o 'generationId=[0-9]*' app.log | tail -200 | sort -u | wc -l
grep 'sending LeaveGroup request' app.log | awk '{print $2}' | sort | uniq -c | sort -rn

Step 2

Find the single late member rather than the average

Break `time-between-poll-max` down per instance. In almost every real case one or two instances sit an order of magnitude above the rest, and they are the ones holding the hot partitions.

kafka-consumer-groups.sh --bootstrap-server broker:9092 --describe --group orders | sort -k6 -n -r | head

Step 3

Check whether the orchestrator is the one killing them

If restarts and rebalances line up, the probe is the cause and Kafka configuration is irrelevant. Look for `OOMKilled`, probe failures, and short container lifetimes before tuning anything.

kubectl get pods -l app=order-consumer -o wide
kubectl describe pod <pod> | sed -n "/Events/,$p"

Step 4

Verify every member agrees on the protocol

The `--members --verbose` output shows the assignor actually in use and the partitions each member holds. A group whose members disagree, or where assignments are wildly uneven, explains both the loop and why one instance is always late.

kafka-consumer-groups.sh --bootstrap-server broker:9092 --describe --group orders --members --verbose

The fix

Break the feedback loop before tuning anything: stop the group, let the backlog be bounded by reducing `max.poll.records` to something the slowest member can certainly finish, then start again. Fixing the arithmetic while the loop is running is guesswork because every measurement is contaminated by the loop itself.

Switch to `CooperativeStickyAssignor`. Incremental cooperative rebalancing revokes only the partitions that must move, so a member joining or leaving no longer stops the entire group. This alone converts most thrashing groups into merely busy ones.

Use static membership for long-lived consumers: set a stable `group.instance.id` per pod and a `session.timeout.ms` longer than a normal restart. A member that restarts within the timeout reclaims its own partitions without any rebalance at all, which removes the deploy-induced storms.

Separate the health endpoint from consumer work — serve probes from a dedicated thread, and report readiness from the consumer’s own liveness signal (time since last successful poll) instead of a generic HTTP 200. A probe that fails because the consumer is busy is a probe that manufactures outages.

Match consumer count to partition count. Idle members add rebalance cost and no throughput; if you need more parallelism, add partitions first and consumers second.

How to stop it coming back

  • Graph `generationId` per group as a first-class metric — it is the earliest, clearest signal of rebalance thrashing
  • Keep per-instance poll metrics, not group averages; this failure is almost always one outlier member
  • Standardise consumer configuration in one shared module so a partial deploy cannot produce protocol disagreement
  • Give every consumer deployment static membership by default; the cost is one environment variable
  • Rehearse a backlog replay in staging before it happens in production — the replay, not steady state, is what breaks the poll interval

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Why does adding consumers make it worse?

Because every member participates in every rebalance, and the join phase waits for all of them. More members means longer rebalances, and longer rebalances mean bigger backlogs when processing resumes, which is what triggers the next rebalance. Until the loop is broken, scaling out feeds it.

What is a normal rebalance duration?

For a healthy group with cooperative assignment, sub-second to a few seconds. Anything in tens of seconds means the join phase is waiting on slow members, usually because they are mid-batch, and that is worth fixing independently of the loop.

Does static membership hide dead consumers?

It delays detection by design: a member that dies is not replaced until `session.timeout.ms` elapses. That is the trade you are making — fewer rebalances in exchange for a longer window before partitions are reassigned. Pick the timeout from how long a normal rolling restart actually takes, plus margin.

Could the brokers be at fault?

Occasionally. A coordinator broker under pressure, or one being moved during maintenance, produces group-wide rebalances that have nothing to do with your handler. The tell is that the `LeaveGroup` reasons are network or coordinator errors rather than `poll timeout has expired`.

Related

Other errors engineers hit next to this one

Full error and symptom index →