Data consistency
Dirty read: interview questions and how to answer them
A dirty read observes another transaction’s uncommitted change, which may later roll back and was never a valid committed state.
Written and reviewed by Sahil Srivastav
What it actually is
Transaction A updates an account but has not committed. Transaction B reads the new balance and makes a decision. A then rolls back, leaving B’s decision based on a value that never existed durably.
Read uncommitted isolation permits this in systems that implement it literally; most production defaults prevent dirty reads, but application caches and replica pipelines can create similar “uncommitted” assumptions at system boundaries.
Why it matters in production
A dirty read can trigger a shipment, credit decision, or alert from a state that disappears milliseconds later. It is more severe than merely reading stale committed data because there is no committed version to reconcile against.
Preventing it does not solve every consistency issue: read committed still allows non-repeatable reads and phantoms.
How it works
Uncommitted visibility
The reader bypasses the writer’s commit boundary and sees a value that may be undone.
Rollback
When the writer aborts, the reader’s observation has no valid database history, but its external side effect may remain.
MVCC prevention
MVCC readers select an older committed version rather than the in-progress version.
Boundary analogue
An event published before its transaction commits is the distributed equivalent: consumers act on a state that may roll back.
Detailed boundary
uncommitted visibility
Operational consequence
why read-uncommitted differs by engine
Implementing it
Use read committed or stronger isolation for business decisions, and never publish events before the local commit.
Keep external side effects behind an outbox and make them idempotent.
Test rollback interleavings, not only successful commits.
Use a two-sided test for this boundary: drive the normal path and the failure path concurrently, then inspect the state that survives the race. For dirty read, the useful assertion is the invariant after recovery, not merely a successful response from one caller.
Document the limit and the signal that tells an operator to change it. A production review of dirty read should name the protected resource, the caller deadline, the expected overload decision, and the evidence that would distinguish a local bug from downstream saturation.
A focused review of dirty read should separate the mechanism from its policy. Reproduce one normal request, one boundary case, and one concurrent failure; record the state transition, the resource consumed, and the signal an operator would see. Then state what the caller is allowed to retry and what must be reconciled manually. This makes dirty read testable in a repository rather than a vocabulary answer.
Interview questions and how to answer them
Give a dirty-read timeline.
A writes a value, B reads it before A commits, then A rolls back. B acted on data that was never committed.
How does MVCC avoid it?
The reader chooses a committed version visible to its snapshot and ignores the in-progress version.
Is read committed enough for every report?
It prevents dirty reads but different statements can see different committed states; a coherent report may need a stable snapshot.
Why is early event publication dangerous?
A consumer can act on an event whose database transaction later rolls back. The outbox couples publication to a committed intent.
What evidence would you inspect for dirty read?
Measure the boundary named in the design, compare it with the caller deadline and resource budget, and reproduce the contention or failure with more than one concurrent worker.
What is the tempting fix for this problem?
Changing a timeout, pool, or retry count alone usually moves the queue. First establish the invariant, then make the bounded mechanism and its failure outcome explicit.
Answers that lose the round
- Using read uncommitted to make reports faster
- Confusing stale committed data with dirty data
- Publishing from inside a transaction before commit
- Assuming a cache transaction rolls back with the database
- Treating “it was eventually corrected” as acceptable for irreversible side effects
- Treating the local mechanism as a complete production guarantee
- Changing the limit without measuring the resource it protects
FAQ
Is dirty read the same as stale read?
No. Stale data was committed but old; dirty data was never committed and may disappear.
Does a lock always prevent dirty reads?
A normal database isolation protocol prevents observing uncommitted writes; explicit lock use alone is not a substitute for the configured visibility rules.
Can read uncommitted ever be safe?
Only for approximate diagnostics where incorrect transient values have no business consequence.