1
0

Add the broker-comparison module

Kafka, RabbitMQ and Pulsar measured side by side on ordering, replay and
consumer scaling by driving the three client libraries directly, plus an
operational-footprint measurement of each broker's own distribution.

Nine tests, three brokers started without Docker, and every number in the
documentation regenerated by scripts/run-all.sh.
This commit is contained in:
2026-09-01 23:27:08 +05:30
parent f4b0112e2c
commit e7e64a047b
30 changed files with 1836 additions and 7 deletions

View File

@@ -0,0 +1,80 @@
prev: [Replay](03-replay.md) · [README](../README.md) · next: [Operational footprint](05-operational-footprint.md)
# 4. Consumer scaling
The question is what happens when you add the fifth consumer.
## Kafka: partitions are a hard ceiling
Five consumers in one group on a three-partition topic
([`docs/output/kafka-consumer-scaling.txt`](output/kafka-consumer-scaling.txt)):
```
consumer-1 -> partitions [0]
consumer-2 -> partitions [1]
consumer-3 -> partitions [2]
consumer-4 -> no partitions (idle)
consumer-5 -> no partitions (idle)
```
A partition belongs to at most one consumer in a group, so partition count is the ceiling. Two of
the five processes are running, connected, healthy, and doing nothing at all — and they will keep
reporting healthy forever.
Raising the partition count later is possible and is not free: it changes which partition a key
maps to, so the per-key ordering guarantee is broken across the change for every key that moves.
In practice the partition count is a capacity decision made at design time, on incomplete
information, that you then live with.
## RabbitMQ: no ceiling, but prefetch decides whether it is real
Five consumers on one queue, forty messages
([`docs/output/rabbit-consumer-scaling.txt`](output/rabbit-consumer-scaling.txt)):
```
no basicQos at all (unlimited prefetch, the AMQP default):
consumer-1 -> 40 consumer-2..5 -> 0
basicQos(1):
consumer-1..5 -> 8 each
```
There is no structural ceiling — but with the AMQP default the broker pushes as many messages as
a consumer will accept, so the first consumer to connect can be handed the entire backlog while
four idle processes wait. `basicQos` is not a tuning knob to postpone; it is what makes the
fan-out exist.
Spring AMQP sets it for you: `AbstractMessageListenerContainer.DEFAULT_PREFETCH_COUNT` is 250.
Better than unlimited, and still enough to concentrate any backlog smaller than 250 messages on
one consumer.
## Pulsar: no ceiling either, and the same trap under a different name
Five consumers, one *non-partitioned* topic, forty messages
([`docs/output/pulsar-consumer-scaling.txt`](output/pulsar-consumer-scaling.txt)):
```
receiverQueueSize default (1000):
one or two consumers take everything; the rest receive nothing
receiverQueueSize(1):
consumer-1..5 -> 8 each
```
The topic has no partitions and five consumers still share the work, which in Kafka would require
at least five partitions decided in advance. But the default receiver queue is a thousand
messages, so the first consumers to connect pull the whole backlog into their own buffers first.
Same failure as RabbitMQ's unlimited prefetch, four times the default.
## The shape of the answer
- Kafka gives you ordering per key and pays for it with a fixed parallelism ceiling.
- RabbitMQ gives you unlimited competing consumers and pays for it with no ordering across them.
- Pulsar lets each subscription choose, which is genuinely the best of both — at the operational
cost measured in [the next chapter](05-operational-footprint.md).
And all three brokers have a client-side buffer that will quietly defeat your fan-out if you leave
it at the default and your backlog is small. That is the one finding here that applies to
whichever you pick.
next: [Operational footprint](05-operational-footprint.md)