Distributed systems

Message queue vs event stream

Distributed systemsMessagingDecision guide

Short answer

Use a message queue when each job should be handled by one competing worker. Use an event stream when multiple consumers need an ordered, retained history they can replay independently. A stream can implement work distribution with consumer groups, but retention and offset operations become part of your design.

Written and reviewed by Sahil Srivastav

What each one actually is

A queue holds messages until a consumer acknowledges them; competing consumers divide work. The message is usually removed or hidden after acknowledgement, making it a natural job buffer.

An event stream appends records to partitions and retains them for a policy or until storage limits. Consumers track offsets, so several applications can read the same event at different speeds or replay it.

The distinction is consumption semantics, not vendor branding. A queue answers “who does this job?”; a stream answers “what happened, and where has each reader reached?”

Side by side

 Message queueEvent stream
Primary purposeDistribute workRetain and distribute facts
ConsumptionOne competing consumer usually handles a messageMany consumer groups read the same record
RetentionUntil ack, timeout, or dead letterIndependent time or size policy
ReplayUsually custom or limitedMove offsets and read again
OrderingOften queue-wide or best-effortGuaranteed only within a partition/key
Backlog meaningUnfinished workReader lag from the retained log
ScalingAdd workers within queue limitsPartitions bound parallelism and ordering
Failure handlingAck timeout, retry, dead letterOffset commit, retry topic, poison record policy

Choose Message queue when

  • Each task should run once by one worker
  • Acknowledge-on-completion and dead-letter handling are sufficient
  • The payload is a command or job rather than an historical fact
  • Consumers do not need independent replay

Choose Event stream when

  • Several applications need the same business event
  • A new consumer must rebuild state from history
  • Per-key ordering and consumer lag are operational signals
  • Readers have different speeds and retention must outlive processing

The trade-off in detail

A stream does not magically make a handler idempotent. A consumer can process a record and crash before committing its offset, so the same record returns after restart.

A queue’s destructive acknowledgement is convenient but makes recovery and audit dependent on dead-letter retention. Preserve the original message, attempt count, and failure reason so operators can repair rather than discard.

Partition keys are a business decision. Choosing customer ID preserves customer order but can create hot partitions; choosing a random key improves spread while losing per-customer sequencing.

Things that are commonly said and are wrong

  • “A queue cannot have multiple consumers.” Competing consumers are normal; the key difference is whether one group shares work or independent groups each see the record.
  • “A stream guarantees global order.” Ordering is normally per partition only.
  • “Acknowledged means business effect committed.” The consumer can acknowledge too early or commit after an external side effect; align processing and state updates deliberately.

Decide it in a real repository

Choosing correctly on a whiteboard and enforcing the choice in code are different skills. Gronex ships broken backend repositories whose tests assert the invariant, not the happy path.

FAQ

Should email delivery use a queue or stream?

A queue is usually the direct fit because one worker should send each email. A stream can be useful when email is one of several projections of a durable business event.

How do I retry a poison message?

Bound attempts, use delayed retries or a retry topic/queue, record the exception and message ID, then dead-letter it with enough context to repair. Infinite immediate retries block healthy work.

Can a stream replace a queue?

Sometimes, if consumer groups, retention, partitioning, and offset management fit the workload. It may add operational cost for a simple job buffer.

Other decisions engineers weigh