Data consistency

Two-phase commit: interview questions and how to answer them

Two-phase commit coordinates several resource managers through prepare and commit phases, preserving atomic decision at the cost of blocking and coordinator dependence.

Written and reviewed by Sahil Srivastav

Data consistencyBackend engineeringInterview preparation

What it actually is

In prepare, each participant performs checks, records a recoverable promise, and votes yes or no. If all vote yes, the coordinator records commit and tells every participant; any no produces rollback.

After voting yes, a participant cannot safely decide alone: the coordinator may have committed while its response was lost. That uncertainty is the source of the blocking failure mode.

Why it matters in production

2PC can keep two databases from observing only half of a transfer, but it holds locks and prepared state while a coordinator or network is unavailable. Under load, a small coordination outage becomes resource exhaustion.

Most service architectures prefer a local transaction plus an outbox or saga because availability and operational simplicity matter more than a single global commit.

How it works

Prepare

Participants validate and persist enough state to commit later, then vote. A yes is a durable promise, not a tentative answer.

Commit decision

The coordinator must durably record the global decision before notifying participants, otherwise recovery cannot distinguish commit from abort.

Blocking

A prepared participant that loses the coordinator waits because unilateral commit or rollback could violate atomicity.

Recovery

Coordinator and participant logs let restarted processes recover the decision, but only if logs and coordinator identity remain available.

Detailed boundary

prepare records

Operational consequence

coordinator failure and blocking

Implementing it

Use 2PC only when the resources support it and the business truly requires atomic cross-resource commit.

Set prepared-transaction alerts and recovery runbooks; a stuck prepared transaction is a correctness and capacity incident.

Prefer an outbox and idempotent consumers when eventual consistency with reconciliation is acceptable.

Use a two-sided test for this boundary: drive the normal path and the failure path concurrently, then inspect the state that survives the race. For two phase commit, the useful assertion is the invariant after recovery, not merely a successful response from one caller.

Document the limit and the signal that tells an operator to change it. A production review of two phase commit should name the protected resource, the caller deadline, the expected overload decision, and the evidence that would distinguish a local bug from downstream saturation.

A focused review of two phase commit should separate the mechanism from its policy. Reproduce one normal request, one boundary case, and one concurrent failure; record the state transition, the resource consumed, and the signal an operator would see. Then state what the caller is allowed to retry and what must be reconciled manually. This makes two phase commit testable in a repository rather than a vocabulary answer.

Interview questions and how to answer them

Why are there two phases?

Participants first promise they can commit, then all apply the same durable decision. Separating the vote from the decision prevents one participant from committing while another cannot.

What happens if the coordinator dies after prepare?

Participants that voted yes may block until the decision is recovered or an external recovery protocol resolves it.

Why choose a saga instead?

A saga avoids global locks by accepting intermediate states and defining compensating actions, improving availability when temporary inconsistency is tolerable.

Can a message queue be a 2PC participant?

Only if it implements a compatible transaction protocol and participates in the same atomic boundary; an ordinary publish API does not.

What evidence would you inspect for two phase commit?

Measure the boundary named in the design, compare it with the caller deadline and resource budget, and reproduce the contention or failure with more than one concurrent worker.

What is the tempting fix for this problem?

Changing a timeout, pool, or retry count alone usually moves the queue. First establish the invariant, then make the bounded mechanism and its failure outcome explicit.

Answers that lose the round

  • Calling 2PC non-blocking
  • Assuming a timeout tells a participant whether to commit
  • Holding locks while the coordinator is unavailable
  • Using it across a provider that cannot participate
  • Ignoring prepared-state recovery
  • Replacing a simple local transaction with distributed coordination
  • Treating the local mechanism as a complete production guarantee
  • Changing the limit without measuring the resource it protects

Practise in a real repository

Explaining a concept and enforcing it in code are different skills, and machine coding rounds test the second. Gronex ships broken backend repositories whose test suites assert the invariant rather than the happy path.

FAQ

Does 2PC guarantee durability?

Only with durable logs and correct participant storage; coordination does not fix unsafe disks.

Is 3PC a universal fix?

It changes failure assumptions but still needs timing and network guarantees that real asynchronous systems may not provide.

What is a prepared transaction?

A participant has persisted a yes vote and is holding resources pending the global decision.

Related

More backend concepts