Data consistency

Saga pattern: interview questions and how to answer them

A saga splits a distributed business transaction into local commits and compensating actions, making progress without holding a global lock.

Written and reviewed by Sahil Srivastav

Data consistencyBackend engineeringInterview preparation

What it actually is

Each step commits in its own service and emits the event that starts the next step. If a later step fails, compensating actions undo the business effect where possible; they do not rewind history or erase the fact that the step happened.

Orchestration centralises the state machine and timeout policy. Choreography lets events drive transitions but can hide the workflow and create cycles. Both require durable state and idempotent handlers.

Why it matters in production

A booking workflow can reserve inventory, take payment, and arrange fulfilment without a transaction spanning three databases. The user sees pending and failed states instead of a coordinator holding locks across network calls.

Compensation is a business operation: a refund may be delayed or rejected, and an email cannot be unsent. The workflow must expose these states and support reconciliation.

How it works

Local commit

Each step commits its own state and an event atomically, commonly through an outbox.

Compensation

The undo action has its own failure modes and may require retries, manual review, or a different compensating value.

Workflow state

Persist step status, attempt count, deadlines, and correlation id so a restart can resume rather than duplicate.

Idempotency

Every command and event handler must tolerate redelivery; the saga cannot assume exactly-once messaging.

Detailed boundary

persisted workflow state

Operational consequence

compensation that can fail

Implementing it

Model pending, completed, compensated, and needs-attention states explicitly.

Use an outbox per local transaction and a deduplicating consumer.

Give each step a timeout and a reconciliation query; a workflow that only advances on happy-path callbacks will stall silently.

Use a two-sided test for this boundary: drive the normal path and the failure path concurrently, then inspect the state that survives the race. For saga pattern, the useful assertion is the invariant after recovery, not merely a successful response from one caller.

Document the limit and the signal that tells an operator to change it. A production review of saga pattern should name the protected resource, the caller deadline, the expected overload decision, and the evidence that would distinguish a local bug from downstream saturation.

A focused review of saga pattern should separate the mechanism from its policy. Reproduce one normal request, one boundary case, and one concurrent failure; record the state transition, the resource consumed, and the signal an operator would see. Then state what the caller is allowed to retry and what must be reconciled manually. This makes saga pattern testable in a repository rather than a vocabulary answer.

Interview questions and how to answer them

How is a saga different from 2PC?

A saga commits local steps independently and compensates on failure; 2PC coordinates one atomic decision but can block resources.

What if compensation fails?

Persist the failure, retry with backoff, and route to reconciliation or manual handling. Hiding it makes the business state lie.

Orchestration or choreography?

Orchestration makes sequencing and timeouts visible; choreography reduces a central coordinator but needs strong event contracts and observability.

How do you prevent duplicate steps?

Use a stable saga and step id with a unique constraint or guarded transition, then return the prior result on replay.

What evidence would you inspect for saga pattern?

Measure the boundary named in the design, compare it with the caller deadline and resource budget, and reproduce the contention or failure with more than one concurrent worker.

What is the tempting fix for this problem?

Changing a timeout, pool, or retry count alone usually moves the queue. First establish the invariant, then make the bounded mechanism and its failure outcome explicit.

Answers that lose the round

  • Calling a compensating action a database rollback
  • Assuming every side effect is reversible
  • Using choreography with no visible workflow state
  • Publishing events outside the local transaction
  • Ignoring duplicate and out-of-order events
  • Returning success before the required business state is reached
  • Treating the local mechanism as a complete production guarantee
  • Changing the limit without measuring the resource it protects

Practise in a real repository

Explaining a concept and enforcing it in code are different skills, and machine coding rounds test the second. Gronex ships broken backend repositories whose test suites assert the invariant rather than the happy path.

FAQ

Are sagas eventually consistent?

Usually. They expose intermediate states and converge through successful steps or compensation.

Can a saga guarantee no overselling?

Only if the inventory step itself enforces capacity atomically; the saga coordinates it but does not replace its invariant.

Do sagas need a message broker?

Not necessarily, but durable asynchronous handoff is common and makes retries and recovery explicit.

Related

More backend concepts