Data consistency
Write-ahead logging: interview questions and how to answer them
Write-ahead logging records a change in durable log storage before the corresponding data page is flushed, allowing crash recovery to replay or undo work.
Written and reviewed by Sahil Srivastav
What it actually is
WAL separates the order of durability from the order of data-page writes. A transaction can commit after its log records are durable even if dirty table pages remain in memory; recovery replays the log after a crash.
The log is an ordered history of changes, but it is not automatically a backup. Retention, archive integrity, and restore testing decide whether it can recover beyond the latest data files.
Why it matters in production
WAL makes group commit and sequential append possible while protecting committed changes from a process or machine crash. It also powers replication and point-in-time recovery.
A full or slow WAL disk stalls commits, and an asynchronous replica may acknowledge less durability than the primary. Storage monitoring is part of correctness.
How it works
WAL rule
Before a dirty data page reaches durable storage, the log describing its change must be durable. Otherwise a crash can leave a page that recovery cannot explain.
Commit record
The commit record is the durability boundary for the transaction. A lost client response does not tell you whether it reached durable log.
Redo
Recovery scans the log and reapplies changes whose data pages were not flushed before the crash.
Checkpoints
A checkpoint records a recovery point and flushes or marks progress, reducing how much log recovery must scan.
Detailed boundary
log-before-data ordering
Operational consequence
recovery and archive retention
Implementing it
Alert on WAL/archive growth, flush latency, replication replay lag, and oldest required segment.
Test restore and point-in-time recovery regularly; a log that cannot be replayed is not a recovery plan.
Choose synchronous replication and commit settings based on the loss window the product accepts.
Use a two-sided test for this boundary: drive the normal path and the failure path concurrently, then inspect the state that survives the race. For write ahead logging, the useful assertion is the invariant after recovery, not merely a successful response from one caller.
Document the limit and the signal that tells an operator to change it. A production review of write ahead logging should name the protected resource, the caller deadline, the expected overload decision, and the evidence that would distinguish a local bug from downstream saturation.
A focused review of write ahead logging should separate the mechanism from its policy. Reproduce one normal request, one boundary case, and one concurrent failure; record the state transition, the resource consumed, and the signal an operator would see. Then state what the caller is allowed to retry and what must be reconciled manually. This makes write ahead logging testable in a repository rather than a vocabulary answer.
Interview questions and how to answer them
Why must the log be flushed before the page?
Recovery needs a durable description of every durable data change. The ordering lets pages be written lazily without losing committed work.
Can a client lose a committed write?
The response can be lost after commit. The client must query by idempotency key or operation id rather than blindly repeat a non-idempotent action.
Is WAL a backup?
It is an ingredient for crash recovery and point-in-time restore; you still need retained base backups and verified archives.
What does replication lag mean here?
The replica has not replayed all primary log records, so reads may be stale even though the primary has committed.
What evidence would you inspect for write ahead logging?
Measure the boundary named in the design, compare it with the caller deadline and resource budget, and reproduce the contention or failure with more than one concurrent worker.
What is the tempting fix for this problem?
Changing a timeout, pool, or retry count alone usually moves the queue. First establish the invariant, then make the bounded mechanism and its failure outcome explicit.
Answers that lose the round
- Treating WAL as a backup by itself
- Assuming commit means every replica applied the row
- Deleting archived segments before validating restore
- Ignoring a full WAL disk
- Confusing log ordering with serialisable isolation
- Running recovery tests only after an incident
- Treating the local mechanism as a complete production guarantee
- Changing the limit without measuring the resource it protects
FAQ
Does WAL prevent data corruption?
It provides a recovery mechanism under its failure model; hardware, bugs, and untested archives can still defeat recovery.
Why are log writes sequential?
Append-oriented writes reduce random I/O and make durable ordering easier to manage.
What happens when WAL fills?
The database may stop accepting writes because required log cannot be recycled or archived safely.