Data consistency

Serializable isolation: interview questions and how to answer them

Serializable isolation makes concurrent execution equivalent to some serial order, using locks, validation, or aborts to prevent anomalies such as write skew.

Written and reviewed by Sahil Srivastav

Data consistencyBackend engineeringInterview preparation

What it actually is

The guarantee is about the final result and permitted observations, not about transactions literally running one at a time. An engine may run them concurrently and then block or abort a transaction when the dependency graph would be cyclic.

Serializable is a correctness tool, not a replacement for schema design. A missing unique constraint or a transaction that calls an external service still needs a separate solution.

Why it matters in production

It prevents subtle anomalies when a rule spans several rows, such as “at least one doctor remains on call”. Two transactions can each see another doctor and independently go off call under weaker isolation; serialisable rejects one.

The price is throughput under contention: retries, lock waits, and long transactions can create a feedback loop.

How it works

Conflict graph

Read/write dependencies form a graph. A cycle means no serial order can explain the execution, so one transaction must wait or abort.

Predicate scope

The rule may concern rows that do not yet exist. Range or predicate protection is needed; locking only existing rows can miss a phantom.

Retry boundary

A serialisation failure invalidates the transaction’s snapshot. Retry the complete transaction, not just its last statement.

Contention control

Short transactions, indexed predicates, and deterministic access order reduce the abort and wait rate.

Detailed boundary

dependency-graph detection

Operational consequence

retrying a complete transaction

Implementing it

Make serialisable retries bounded and attach an idempotency key to the outer request.

Keep external calls outside the transaction and record an intent first.

Load test the invariant with conflicting writers; average latency hides abort storms.

Use a two-sided test for this boundary: drive the normal path and the failure path concurrently, then inspect the state that survives the race. For serializable isolation, the useful assertion is the invariant after recovery, not merely a successful response from one caller.

Document the limit and the signal that tells an operator to change it. A production review of serializable isolation should name the protected resource, the caller deadline, the expected overload decision, and the evidence that would distinguish a local bug from downstream saturation.

A focused review of serializable isolation should separate the mechanism from its policy. Reproduce one normal request, one boundary case, and one concurrent failure; record the state transition, the resource consumed, and the signal an operator would see. Then state what the caller is allowed to retry and what must be reconciled manually. This makes serializable isolation testable in a repository rather than a vocabulary answer.

Interview questions and how to answer them

Why can a serialisable transaction fail after all statements succeeded?

The engine may discover at commit that another transaction created a dependency cycle. Aborting preserves the serialisable result.

What is write skew?

Two transactions read a shared rule, update different rows, and jointly violate the rule. Serializable detects the cross-row dependency; row locks on only one row may not.

What should a retry include?

Begin a new transaction, reread all inputs, recompute, and reapply. Reusing the old snapshot is incorrect.

When is explicit locking preferable?

When the invariant is narrow and a deterministic row or range lock can protect it with lower abort overhead.

What evidence would you inspect for serializable isolation?

Measure the boundary named in the design, compare it with the caller deadline and resource budget, and reproduce the contention or failure with more than one concurrent worker.

What is the tempting fix for this problem?

Changing a timeout, pool, or retry count alone usually moves the queue. First establish the invariant, then make the bounded mechanism and its failure outcome explicit.

Answers that lose the round

  • Assuming serialisable means no aborts
  • Retrying only the failed UPDATE
  • Holding a serialisable transaction across network I/O
  • Forgetting phantom rows
  • Using it to compensate for missing constraints
  • Reporting every abort as a server error instead of a bounded retry or conflict
  • Treating the local mechanism as a complete production guarantee
  • Changing the limit without measuring the resource it protects

Practise in a real repository

Explaining a concept and enforcing it in code are different skills, and machine coding rounds test the second. Gronex ships broken backend repositories whose test suites assert the invariant rather than the happy path.

FAQ

Does serialisable prevent deadlocks?

No. Lock-based implementations can still deadlock and need retries.

Is serialisable stronger than repeatable read?

Yes for the anomalies it prevents, especially cross-row write skew and phantoms, though exact behaviour is engine-specific.

Can serialisable span two databases?

No. A local database transaction cannot atomically coordinate another resource without a distributed protocol.

Related

More backend concepts