Lesson 2

Who reads what

Inside a consumer group each partition goes to exactly one consumer. An assignor class decides the split, and Kafka ships four of them that produce four different answers for the same configuration. This lesson measures all four on Kafka 4.3.1, including a case where nine partitions shared among four consumers still leave one consumer with nothing.

The only rule

A partition goes to at most one consumer in the same group. Two consequences follow, and both are widely known: the number of working consumers never exceeds the number of partitions, and the partition count is a hard ceiling on parallelism. What gets less attention is the rest — how the split is made — which is not a detail. It decides whether that ceiling is actually reached.

The measurements below build groups with the tools that ship with Kafka and read the result back:

kafka-console-consumer.sh --include 't3[abc]' --group g1 \
  --consumer-property group.protocol=classic \
  --consumer-property partition.assignment.strategy=\
org.apache.kafka.clients.consumer.RangeAssignor

kafka-consumer-groups.sh --describe --group g1 --members --verbose
# CLIENT-ID  #PARTITIONS  CURRENT-ASSIGNMENT
# c1         6            t3a:0,1;t3b:0,1;t3c:0,1
# c2         3            t3a:2;t3b:2;t3c:2

range splits per topic

RangeAssignor is the default. It handles each topic independently: per topic, every consumer takes ⌊P/C⌋ consecutive partitions, and the first P mod C consumers take one more.

On a single topic that is reasonable. Across several topics the same split repeats verbatim, so every topic's remainder lands on the same consumers at the head of the list. Three topics each with one partition left over means the first consumer collects three extras; the remainder is not spread around.

Measured on three 3-partition topics with two consumers:

consumerassignmentpartitions
c1t3a:0,1;t3b:0,1;t3c:0,16
c2t3a:2;t3b:2;t3c:23

Six against three: one consumer carries twice the load. On the same configuration roundrobin gives 5 and 4.

The case worth remembering: 9 partitions, 4 consumers, one idle Still three 3-partition topics, now four consumers. Each topic splits 3 partitions among 4 consumers, so ⌊3/4⌋ = 0 and the first three consumers take one each. The fourth takes nothing — in every topic. Measured: 3, 3, 3, 0. Roundrobin on the same configuration: 3, 2, 2, 2.

roundrobin spreads the group

RoundRobinAssignor collects every topic–partition pair in the group into one list, sorted by topic name then partition number, and deals them round the table. Because the list is never cut along topic boundaries, one topic's remainder is offset by the next topic's.

Same three 3-partition topics, two consumers:

consumerassignmentpartitions
c1t3a:0,2;t3b:1;t3c:0,25
c2t3a:1;t3b:0,2;t3c:14

Any two consumers differ by at most one partition, which is simply what dealing round the table gives you. With ten consumers over 9 partitions, roundrobin produced nine consumers with one partition each and one idle, while range produced three consumers with three partitions each and seven idle. Same group, same topics.

Uneven topics show the difference just as clearly: one 5-partition topic plus one 1-partition topic, two consumers. Range gives 4 and 2; roundrobin gives 3 and 3.

sticky and cooperative-sticky

StickyAssignor balances like roundrobin but arrives there differently: it orders partitions by index across topics — every partition 0 first, then every partition 1 — and then fills each consumer up to its quota instead of dealing round the table. On three 3-partition topics with two consumers:

consumerassignmentpartitions
c1t3a:0,1;t3b:0,1;t3c:05
c2t3a:2;t3b:2;t3c:1,24

Still 5 and 4, grouped differently: exactly the interleaved list t3a:0, t3b:0, t3c:0, t3a:1, t3b:1, t3c:1, t3a:2, t3b:2, t3c:2 cut after the fifth element.

That interleaved rule only holds up to three consumers It matches the broker partition for partition at 2 and 3 consumers, and at 2 consumers over the uneven 5 + 1 pair. At 4 consumers, however, the measured shares were t3a:0;t3b:0;t3c:0, t3a:1;t3b:1, t3a:2;t3b:2 and t3c:1,2: the sizes are still 3, 2, 2, 2 as the quota predicts, but the grouping is different. No static ordering explains all three group sizes, so with sticky rely on the share sizes and nothing finer. The lab below states that limit right under its result.

The name “sticky” is about later assignments: when a consumer leaves, it tries to keep the remaining consumers on what they already had rather than reshuffling everything. The lab here models only the first assignment, when there is nothing to stick to.

cooperative-sticky splits identically to sticky Run both on the same three topics with two consumers: the results match partition for partition. The difference is not in the split but in the rebalance protocol — cooperative revokes only what genuinely has to move, while eager makes the whole group drop everything and re-acquire. For a large group that is the difference between a short pause and a long one, but the final assignment table is the same.

The numbers

Three topics, 3 partitions each, 9 partitions in total. Partitions received per consumer:

consumersrangeroundrobinstickyidle
26, 35, 45, 4none
33, 3, 33, 3, 33, 3, 3none
43, 3, 3, 03, 2, 2, 23, 2, 2, 2range: 1
103, 3, 3 and seven zerosnine ones, one zeronine ones, one zerorange: 7

Uneven topics — one with 5 partitions, one with 1, two consumers:

strategyc1c2
rangeu1:0;u5:0,1,2 — 4 partitionsu5:3,4 — 2 partitions
roundrobinu1:0;u5:1,3 — 3 partitionsu5:0,2,4 — 3 partitions
stickyu1:0;u5:0,1 — 3 partitionsu5:2,3,4 — 3 partitions

With a single topic the three strategies are equivalent in count and differ only in which partition goes where: on one 3-partition topic with two consumers, range gives 0,1 and 2 while roundrobin gives 0,2 and 1. In other words, every difference in this lesson only appears once a group reads from more than one topic. For a single-topic group the choice does not matter.

What is and is not predictable

The lab's engine reproduces the set of shares in all 16 measured scenarios. One thing it does not reproduce, and should not promise: which consumer gets which share.

Ten consumers named c1…c10 were run into a range group reading one 3-partition topic twice, differing only in the order they joined. The first run gave the partitions to c1, c10, c2; the second to c1, c2, c3. A rerun of roundrobin likewise produced a different set from the run before. The internal ordering cannot be derived from client ids.

Do not design around identity A system that relies on “consumer 1 always reads partition 0” will pass a test and break after the first rebalance. What is predictable is the shape of the split — how many partitions each consumer holds, and whether anyone is left with none — and that is also the only part worth sizing against.

The lab

Set your own topics, partition counts and consumer count, then switch strategies to compare. The algorithms are rewrites of the three assignors and reproduce the share sets of all 16 scenarios measured on Kafka 4.3.1.

Who reads what

    Configuration

    Takeaways

    The previous lesson covers the other half: which partition a key lands on.