Data consistency

Read committed vs repeatable read: interview questions and how to answer them

Read committed gives each statement a current committed view, while repeatable read keeps a transaction-level snapshot or equivalent guarantee; neither automatically fixes every write race.

Written and reviewed by Sahil Srivastav

Data consistencyBackend engineeringInterview preparation

What it actually is

Under read committed, a second SELECT can see a commit that happened after the first SELECT. Under PostgreSQL repeatable read, the transaction reads one snapshot and a concurrent conflicting update can cause a serialisation failure.

The names are not portable descriptions of every detail. InnoDB’s repeatable read includes next-key locking for some writes, so the same SQL can block differently across engines.

Why it matters in production

Read committed is a good default for independent statements, but a report assembled from several statements can mix points in time. Repeatable read gives a coherent view for that report.

A stable snapshot does not mean the transaction may safely overwrite what it read. The write predicate, lock, or version must still express ownership of the decision.

How it works

Statement snapshot

Read committed selects a visibility point for each statement, so later statements can see newer commits.

Transaction snapshot

Repeatable read holds one visibility point for reads, producing a coherent view but potentially retaining old row versions.

Concurrent writes

A transaction can read an old value and later fail or overwrite depending on engine semantics and the UPDATE predicate.

Operational cost

Long snapshots delay vacuum or retain versions; lock waits and aborted transactions are part of the level’s cost model.

Detailed boundary

statement versus transaction snapshots

Operational consequence

PostgreSQL snapshot isolation

Implementing it

Use read committed for short independent changes with atomic predicates. Use a stable snapshot for a multi-query report that must be internally coherent.

Instrument transaction age, snapshot retention, lock waits, and serialisation failures.

Add a version or conditional update when a user edits data across requests; no isolation level spans the user’s think time.

Use a two-sided test for this boundary: drive the normal path and the failure path concurrently, then inspect the state that survives the race. For read committed vs repeatable read, the useful assertion is the invariant after recovery, not merely a successful response from one caller.

Document the limit and the signal that tells an operator to change it. A production review of read committed vs repeatable read should name the protected resource, the caller deadline, the expected overload decision, and the evidence that would distinguish a local bug from downstream saturation.

A focused review of read committed vs repeatable read should separate the mechanism from its policy. Reproduce one normal request, one boundary case, and one concurrent failure; record the state transition, the resource consumed, and the signal an operator would see. Then state what the caller is allowed to retry and what must be reconciled manually. This makes read committed vs repeatable read testable in a repository rather than a vocabulary answer.

Interview questions and how to answer them

Can two reads in read committed differ?

Yes. A concurrent commit between statements can be visible to the second statement.

What does repeatable read buy?

A coherent read view for the transaction, at the cost of retained versions and possible write conflicts.

Does repeatable read prevent a lost update?

Not as a general application contract. Use a guarded UPDATE, lock, or version column.

Which would you use for a dashboard?

Usually a short consistent snapshot if cross-widget coherence matters; otherwise read committed with an explicit freshness label.

What evidence would you inspect for read committed vs repeatable read?

Measure the boundary named in the design, compare it with the caller deadline and resource budget, and reproduce the contention or failure with more than one concurrent worker.

What is the tempting fix for this problem?

Changing a timeout, pool, or retry count alone usually moves the queue. First establish the invariant, then make the bounded mechanism and its failure outcome explicit.

Answers that lose the round

  • Calling read committed inconsistent by definition
  • Assuming repeatable read makes a detached UI edit safe
  • Ignoring engine-specific semantics
  • Keeping a report transaction open while streaming to a slow client
  • Retrying an aborted transaction with the same stale inputs
  • Treating the local mechanism as a complete production guarantee
  • Changing the limit without measuring the resource it protects

Practise in a real repository

Explaining a concept and enforcing it in code are different skills, and machine coding rounds test the second. Gronex ships broken backend repositories whose test suites assert the invariant rather than the happy path.

FAQ

Why can repeatable read cause bloat?

Old row versions cannot be removed while an old snapshot may still need them.

Does read committed permit dirty reads?

Normally no: uncommitted changes are not visible.

Can read committed be serialisable for one statement?

A single atomic statement can enforce a strong invariant without making the entire transaction serialisable.

Related

More backend concepts