Data consistency
MVCC: interview questions and how to answer them
Multi-version concurrency control lets readers use a consistent snapshot while writers create newer row versions, reducing read–write blocking at the cost of version cleanup.
Written and reviewed by Sahil Srivastav
What it actually is
MVCC stores visibility metadata or row versions so a transaction can choose the version valid for its snapshot. Readers generally do not wait for an uncommitted writer, and writers do not overwrite a version still needed by an active snapshot.
An update is often an insert of a new physical version plus a marker on the old version. Vacuum or garbage collection later reclaims versions no snapshot can see.
Why it matters in production
MVCC keeps ordinary reads available during writes, which is essential for busy APIs and dashboards. It also explains why a forgotten transaction can cause table or undo-log growth even when disk writes look normal.
Snapshot isolation can still permit write skew, and a replica can serve a perfectly consistent but stale snapshot. MVCC is not a global freshness guarantee.
How it works
Visibility
A snapshot records which transactions are visible. A row version created after the snapshot is hidden; an older version remains readable if it was committed in time.
Writer conflict
Two writers targeting the same logical row cannot both commit incompatible versions without one waiting or aborting.
Cleanup
Vacuum or equivalent cleanup removes dead versions after the oldest active snapshot no longer needs them.
Long readers
A long report or idle transaction pins old versions and can make cleanup fall behind, increasing I/O and index bloat.
Detailed boundary
snapshot visibility
Operational consequence
old-version reclamation
Implementing it
Monitor oldest transaction age, dead tuples, undo history, and cleanup lag.
Commit or roll back promptly; never leave a transaction open while waiting for a client.
Use serialisable or explicit predicates when the invariant spans rows; MVCC alone is not enough.
Use a two-sided test for this boundary: drive the normal path and the failure path concurrently, then inspect the state that survives the race. For mvcc, the useful assertion is the invariant after recovery, not merely a successful response from one caller.
Document the limit and the signal that tells an operator to change it. A production review of mvcc should name the protected resource, the caller deadline, the expected overload decision, and the evidence that would distinguish a local bug from downstream saturation.
A focused review of mvcc should separate the mechanism from its policy. Reproduce one normal request, one boundary case, and one concurrent failure; record the state transition, the resource consumed, and the signal an operator would see. Then state what the caller is allowed to retry and what must be reconciled manually. This makes mvcc testable in a repository rather than a vocabulary answer.
Interview questions and how to answer them
How can a reader avoid blocking a writer?
It reads an older committed version visible to its snapshot while the writer creates a new version.
Why does bloat grow?
Dead versions cannot be reclaimed while an active snapshot might still need them.
Does MVCC prevent lost updates?
It detects or permits them according to engine and write pattern; use a guarded update, lock, or version.
What operational metric matters most?
The age of the oldest transaction and cleanup lag, because they predict retained versions and storage pressure.
What evidence would you inspect for mvcc?
Measure the boundary named in the design, compare it with the caller deadline and resource budget, and reproduce the contention or failure with more than one concurrent worker.
What is the tempting fix for this problem?
Changing a timeout, pool, or retry count alone usually moves the queue. First establish the invariant, then make the bounded mechanism and its failure outcome explicit.
Answers that lose the round
- Calling MVCC a lock-free database
- Assuming snapshots are current forever
- Ignoring long idle transactions
- Expecting repeatable reads to prevent write skew
- Running large reports without resource limits
- Treating vacuum as optional maintenance
- Treating the local mechanism as a complete production guarantee
- Changing the limit without measuring the resource it protects
FAQ
Does MVCC mean no locks?
No. Writers, schema changes, indexes, and cleanup still use locks.
Is a snapshot the same as a backup?
No. A snapshot is a transaction visibility view, not necessarily durable or restorable.
Why can a replica be stale with MVCC?
Replication delay determines which committed versions have arrived; local snapshot consistency does not imply primary freshness.