Distributed systems

Kafka vs RabbitMQ

Distributed systemsMessagingDecision guide

Short answer

Choose Kafka when durable event history, replay, and high-throughput partitioned consumption are central. Choose RabbitMQ when routing commands to workers, per-message acknowledgement, and low-latency queue semantics matter more; do not use Kafka as a job queue simply because it is popular.

Written and reviewed by Sahil Srivastav

What each one actually is

Kafka is a distributed append-only log. Producers write records to partitions, consumers track offsets, and retention lets independent groups replay at their own pace.

RabbitMQ is a broker built around exchanges, queues, bindings, acknowledgements, and routing. A message is normally delivered to one competing consumer and removed after acknowledgement.

Both can lose or duplicate effects if the application commits at the wrong point. Delivery guarantees describe transport behaviour, not an end-to-end business transaction.

Side by side

 KafkaRabbitMQ
Core modelPartitioned retained logBrokered queues and exchanges
ReplayMove offsets within retentionUsually republish or dead-letter a message
RoutingTopic and partition keyRich exchange and binding patterns
Work distributionConsumer groupsCompeting consumers with ack
OrderingPer partitionQueue ordering is affected by concurrency and redelivery
Backpressure signalConsumer lagQueue depth and unacked count
Payload lifetimeRetention policyUntil ack, expiry, or dead letter
Operational shapePartition, broker, and offset managementQueue, exchange, connection, and ack management

Choose Kafka when

  • Several independent consumers need the same event
  • A new consumer must rebuild a projection from history
  • Throughput and partitioned parallelism dominate routing flexibility
  • You can monitor lag and manage retention and replays

Choose RabbitMQ when

  • Each command should be processed by one worker
  • Routing rules need exchanges, bindings, and per-message delivery
  • Work should disappear after successful acknowledgement
  • The workload is moderate and queue semantics are more useful than replay

The trade-off in detail

Kafka’s retained log makes recovery powerful, but retention is not a free archive. A consumer that falls behind can exhaust storage or miss the business recovery window; monitor lag and size retention against the slowest rebuild.

RabbitMQ’s acknowledgements make work completion explicit, but unacked messages consume broker and consumer capacity. Ack after the side effect is durable, and cap prefetch so one worker cannot reserve the whole queue.

Ordering and retries interact in both systems. A failed Kafka record can block a partition if strict order is required; RabbitMQ redelivery can put a message back in a surprising position. Choose whether ordering or progress wins for each key.

Things that are commonly said and are wrong

  • “Kafka is just a faster RabbitMQ.” Their storage and consumption models differ; the correct choice follows replay and routing needs.
  • “RabbitMQ cannot scale.” It can cluster and route across nodes, but its scaling model and ordering trade-offs differ from Kafka partitions.
  • “Kafka exactly once makes external effects exactly once.” External databases and APIs still need idempotency or a coordinated commit.

Decide it in a real repository

Choosing correctly on a whiteboard and enforcing the choice in code are different skills. Gronex ships broken backend repositories whose tests assert the invariant, not the happy path.

FAQ

Which should I use for background jobs?

RabbitMQ is often the simpler fit for commands that one worker acknowledges. Kafka is suitable when the job is also a durable event and replay, multiple consumer groups, or very high throughput matters.

Can RabbitMQ replay messages?

Not in the same native offset-log way. Retain or republish messages deliberately, and preserve IDs and attempt metadata so replay does not create duplicate effects.

What should Kafka consumers commit?

Commit the offset only after the processing effect is durable, or use an atomic integration supported by the chosen sink. A commit before the effect loses work; after it can redeliver and therefore requires idempotency.

Other decisions engineers weigh