Distributed systems

Event sourcing: interview questions and how to answer them

Event sourcing stores the sequence of state changes as the source of truth, and derives current state by replaying them.

Written and reviewed by Sahil Srivastav

Distributed systemsArchitectureTrade-offs

What it actually is

In an event-sourced system, you do not store the current state of an entity. You store the ordered sequence of events that produced it — OrderPlaced, ItemAdded, PaymentCaptured, OrderShipped — and current state is a fold over that sequence. The event log is the authoritative record and everything else is derived.

What this buys is history as a first-class property rather than something bolted on. You can answer what the state was at any past moment, why it changed, and in what order — without an audit table that someone has to remember to write to. You can also build a new read model over historical data, because the events needed to populate it were never discarded.

The honest framing is that it is a significant architectural commitment that pays off for a narrow set of domains. Where the history is the business — ledgers, trading, insurance claims, regulated workflows — it fits naturally, because those domains already think in terms of immutable entries. For typical CRUD, it adds substantial complexity in exchange for benefits the product does not need.

Why it matters in production

Because the costs are real and arrive later than the benefits, which is what makes it a frequent regret. Querying is the first one: there is no table to select from, so every query needs a projection, and a question nobody anticipated requires building and backfilling a new one. Teams used to writing an ad-hoc query against current state find this genuinely disruptive.

The deeper cost is that events are immutable and permanent, so every mistake is permanent too. A bug that emitted wrong events cannot be fixed by an UPDATE — you need compensating events or a rewritten stream. Schema evolution has to handle events written years ago by code that no longer exists. And a legal deletion request conflicts directly with an append-only log, which is a problem with no clean answer.

How it works

State is a fold, snapshots are an optimisation

Current state is produced by replaying events in order. Replaying thousands of events per read is too slow, so you periodically persist a snapshot and replay only events after it. The snapshot is strictly derived data and must be safe to delete and rebuild — a snapshot treated as authoritative has quietly abandoned the model.

Optimistic concurrency via expected version

Appends specify the version the writer believed was current. If another writer appended first, the version no longer matches and the append is rejected, so the writer re-reads and retries. This is how an append-only log enforces invariants without locking, and it is the mechanism that makes concurrent command handling safe.

Projections are rebuildable read models

Queries are served by projections built from the event stream and shaped for specific reads. Because the events are retained, a projection can be dropped and rebuilt from scratch — which is the real benefit, since a new question becomes a new projection rather than a data migration.

Schema evolution is the hard, permanent part

Events written years ago cannot be changed, so the code must read every version it ever wrote. The standard approaches are upcasting — transforming old shapes into current ones on read — or versioned event types handled explicitly. Either way, this obligation never goes away and grows with the system's age.

Deletion conflicts with immutability

A right-to-erasure request cannot be satisfied by appending a "deleted" event, because the personal data remains in the log. The practical answer is crypto-shredding: store personal data encrypted with a per-subject key and destroy the key, rendering the events unreadable while leaving the log intact. This needs designing up front — retrofitting it is extremely painful.

Implementing it

Scope it to the aggregates where history genuinely is the business value, rather than adopting it system-wide. Mixing an event-sourced ledger with conventional CRUD elsewhere is normal and usually correct.

Design events as business facts in past tense — PaymentCaptured, not UpdatePaymentStatus. Events describing CRUD operations give up the semantic benefit and leave you with the costs.

Plan schema evolution and personal-data handling before the first event is written. Both become dramatically harder once a production log exists.

Keep projections genuinely disposable and rebuild them regularly in non-production. A projection that cannot be rebuilt has silently become a second source of truth.

// State is derived, never stored directly.
const state = events.reduce(apply, initialState);

// Appending with an expected version: the log enforces the invariant.
await store.append(streamId, newEvents, { expectedVersion: 7 });
// Throws if someone else appended first -> re-read, re-decide, retry.

// Snapshot is an optimisation and must be safe to delete.
const snap = await snapshots.latest(streamId);        // version 500
const tail = await store.read(streamId, { after: snap.version });
const current = tail.reduce(apply, snap.state);

Interview questions and how to answer them

What does event sourcing actually buy you?

Complete history as a structural property rather than an add-on: what the state was at any time, why it changed, in what order, with no separate audit table to maintain. And the ability to build new read models retroactively, because the events needed were never discarded. Both matter most in domains where history is the business — ledgers, claims, regulated workflows.

How do you query current state?

Through a projection built from the stream and shaped for that query. There is no table holding current state to select from. This is the cost people underestimate: a question nobody anticipated requires building a projection and backfilling it, where a conventional schema would have allowed an ad-hoc query.

An event was emitted with wrong data. How do you fix it?

Not with an update — the log is immutable. You append a compensating event that corrects the state going forward, which is also the honest record of what happened. If the stream is genuinely corrupt rather than merely wrong, you rewrite it into a new stream and cut over, which is a serious operation. This permanence is why event design deserves real care.

How do you handle a right-to-erasure request?

Appending a deletion event does not remove the personal data from the log, so it does not satisfy the requirement. The usual answer is crypto-shredding: encrypt personal data with a per-subject key held outside the log, and destroy the key on request. The events remain structurally intact and the personal data becomes unrecoverable. It must be designed in from the start.

When would you not use event sourcing?

For most CRUD. If nobody asks what the state was last Tuesday, the history adds complexity without value, and you pay the full cost — projections for every query, permanent schema evolution, deletion complications — for a benefit the product does not use. The default should be a conventional schema with an audit log where needed.

Answers that lose the round

  • Adopting it system-wide when only one aggregate benefits from the history
  • Naming events as CRUD operations, losing the semantic value while keeping the cost
  • Treating snapshots as authoritative rather than as a rebuildable cache
  • No plan for reading events written by versions of the code that no longer exist
  • Discovering the deletion-versus-immutability conflict after going live
  • Expecting to query current state ad hoc, with no projection for the question
  • Confusing it with CQRS — related, separable, frequently conflated

Practise in a real repository

Explaining a concept and enforcing it in code are different skills, and machine coding rounds test the second. Gronex ships broken backend repositories whose test suites assert the invariant rather than the happy path.

FAQ

Is event sourcing the same as CQRS?

No, though they are often paired. CQRS separates read and write models and can be applied over an ordinary database. Event sourcing is about the source of truth being a log of events. Event sourcing nearly always implies CQRS, because reads need projections, but CQRS does not require event sourcing.

Can I use Kafka as the event store?

Carefully. Kafka is a log but lacks what an event store provides: reading a single entity's stream efficiently, and optimistic concurrency on append. People do it with per-entity partitioning and external version checks, but a purpose-built event store — or an append-only table in a relational database — is usually a better fit for the write side.

How do snapshots affect correctness?

They must not. A snapshot is a cache of a fold, so deleting all snapshots should change nothing but latency. Rebuild them in non-production routinely to verify that, because a snapshot that cannot be regenerated has become an undeclared source of truth.

Does it work with microservices?

It composes well, since published events are a natural integration point. The caution is that internal events and published contracts should be separate: making your internal event shapes a public interface means you can never refactor them without breaking consumers.

Related

More backend concepts