Distributed systems

Synchronous vs asynchronous communication

Distributed systemsMessagingDecision guide

Short answer

Use synchronous communication when the caller needs a bounded answer before it can continue. Use asynchronous communication when work can complete later, must absorb bursts, or should survive a temporary consumer outage; make the resulting status, retries, and duplicate handling explicit.

Written and reviewed by Sahil Srivastav

What each one actually is

A synchronous call holds a caller until the callee returns success or failure. It gives a direct result and simple control flow, but the caller inherits the callee’s latency and availability.

Asynchronous communication records work or an event and lets a consumer process it later. Queues absorb bursts and decouple lifetimes, but the caller needs a receipt, status endpoint, event, or other way to learn the outcome.

“Async” can mean a background task inside one process or a durable message across services. Only durable storage and a recovery policy provide resilience across process failure.

Side by side

 Synchronous communicationAsynchronous communication
Caller latencyIncludes downstream workUsually returns an acknowledgement quickly
Burst handlingConcurrency reaches downstream immediatelyQueue can buffer and apply backpressure
Failure visibilityImmediate responseFailure appears later through status or dead letter
Duplicate handlingRetries can still duplicate writesAt-least-once delivery makes idempotency essential
OrderingCall sequence is visible to callerPartitions and retries can reorder work
CouplingRuntime availability and schema couplingProducer and consumer can be deployed separately
User experienceNatural for validation and readsNatural for import, email, and long-running work
OperationsTimeouts and pool pressureLag, poison messages, replay, and DLQ monitoring

Choose Synchronous communication when

  • The caller cannot proceed without a result
  • Validation or a small read must be shown immediately
  • The dependency is local and its latency budget is bounded
  • A user action must fail visibly before the response ends

Choose Asynchronous communication when

  • Work is slow, bursty, or retryable
  • A consumer may be temporarily unavailable
  • The producer should not hold a connection during processing
  • A durable event or job can be replayed and deduplicated

The trade-off in detail

Async removes waiting from the request but does not remove waiting from the business process. Users need a durable job identifier and a truthful state model such as accepted, processing, succeeded, and failed.

Synchronous retries can amplify an outage because every caller retries at once. Queues shift that pressure to consumers, where bounded concurrency and explicit backpressure are easier to control.

A queue is not a transaction across systems. Use an outbox when a database change and message publication must agree, then make consumers idempotent because redelivery is normal.

Things that are commonly said and are wrong

  • “Asynchronous means faster.” It improves caller latency and elasticity; total completion can be slower.
  • “A successful enqueue means the job succeeded.” It means durable acceptance only.
  • “Synchronous calls are exactly once.” Network failures can occur after the server commits and before the response arrives.

Decide it in a real repository

Choosing correctly on a whiteboard and enforcing the choice in code are different skills. Gronex ships broken backend repositories whose tests assert the invariant, not the happy path.

FAQ

Should checkout be synchronous?

Keep the small decision that must be shown immediately synchronous, but move fulfilment, email, and downstream notifications to durable asynchronous work. Return an order state and make retries safe.

How should an async API report errors?

Persist a status and error category keyed by a job or request ID, expose polling or a webhook, and retain enough detail for support. A dead-letter queue is an operator tool, not a customer-facing state.

Can async messages be exactly once?

Delivery is commonly at least once. You can make the business effect effectively once with an idempotency key, a durable processed marker, and an atomic commit strategy.

Other decisions engineers weigh