Distributed systems

Orchestration vs choreography

Distributed systemsWorkflowsDecision guide

Short answer

Choose orchestration when a business workflow needs one visible owner, explicit branching, and compensations. Choose choreography for simple event reactions owned by independent services; once a workflow has many steps or hidden dependencies, centralising its state is usually safer than making operators reconstruct it from events.

Written and reviewed by Sahil Srivastav

What each one actually is

An orchestrator owns the workflow state and calls participants or sends commands. It can show the next step, deadline, retry policy, and compensation in one place.

In choreography, services subscribe to events and decide their own reactions. There is no central coordinator, so coupling is reduced at the interface but the overall workflow is spread across consumers.

Both are ways to implement a saga: neither provides a distributed ACID transaction. Every step needs an idempotency key and a recovery story.

Side by side

 OrchestrationChoreography
Workflow visibilityOne state machine is inspectableMust correlate events across services
CouplingCoordinator knows participant contractsConsumers depend on event semantics
Adding a stepChange one workflow definitionAdd a subscriber, but discover all side effects
BranchingExplicit conditions and timeoutsEmerges from event reactions
Failure recoveryCentral compensation and retry policyEach service owns recovery for its reaction
Operational riskCoordinator can become a bottleneckHidden cycles and duplicate reactions
Team autonomyCentral workflow ownershipTeams own event consumers independently
TestingState-machine and contract testsEvent contracts plus end-to-end scenarios

Choose Orchestration when

  • A payment, reservation, or fulfilment flow has ordered steps and compensation
  • Operators need one place to see and resume a stuck workflow
  • Timeouts and human approval are explicit parts of the process
  • A single team owns the business process

Choose Choreography when

  • Events are independent notifications with simple reactions
  • Services should react without knowing the full workflow
  • The event contract is stable and consumers can evolve independently
  • Central workflow ownership would be artificial or too tightly coupled

The trade-off in detail

Choreography looks decoupled until nobody knows who emits the event that drives a side effect. Maintain an event catalogue, ownership, correlation IDs, and a way to replay safely; otherwise debugging becomes archaeology.

An orchestrator is not automatically a synchronous bottleneck. It can persist state and issue commands asynchronously, but it becomes a critical dependency and must be made highly available and idempotent.

Compensations are new business actions, not database rollbacks. A refund, release, or cancellation can itself fail, so model it as a state transition with retries and operator visibility.

Things that are commonly said and are wrong

  • “Choreography has no coupling.” Consumers are coupled to event names, fields, ordering, and meaning.
  • “Orchestration means one giant service.” A focused workflow component can own coordination while participants own their data.
  • “A saga guarantees consistency.” It provides a path to convergence through compensating actions, not atomic isolation.

Decide it in a real repository

Choosing correctly on a whiteboard and enforcing the choice in code are different skills. Gronex ships broken backend repositories whose tests assert the invariant, not the happy path.

FAQ

Which is easier to operate?

Orchestration usually is once a workflow has several steps, because state and retries are visible together. Choreography is easy for small independent reactions but needs strong event observability as it grows.

Can a system use both?

Yes. An order saga may be orchestrated, while independent analytics and notifications subscribe to the resulting events. Keep commands and facts distinct so consumers do not accidentally become workflow owners.

Where should saga state live?

In durable storage owned by the orchestrator, keyed by a correlation or business ID. Memory-only coordination loses progress on restart and makes recovery depend on guessing.

Other decisions engineers weigh