Distributed systems
Consensus and Raft: interview questions and how to answer them
Consensus gets a group of nodes to agree on an ordered sequence of operations despite failures, which is what makes replicated state safe.
Written and reviewed by Sahil Srivastav
What it actually is
Consensus is the problem of getting several nodes to agree on a value — in practice, on an ordered log of operations — when nodes can crash and messages can be delayed or lost. Agreement on the order is what lets every replica apply the same operations and arrive at the same state, which is the foundation under replicated databases, configuration stores and lock services.
Raft is a consensus algorithm designed specifically to be understandable, which is why it displaced Paxos in practice and why interviews ask about it. It decomposes the problem into three pieces that can be reasoned about separately: electing a leader, replicating a log from that leader, and guaranteeing safety across leader changes.
The property doing the real work is majority quorum. Any two majorities of the same set must share at least one node, so a decision accepted by one majority cannot be contradicted by another — the overlapping node would have to have accepted both. That single observation is why split brain is impossible and why the answer to "why a majority?" is not "for redundancy".
Why it matters in production
Because it is what stands between a replicated system and silently divergent state. Without it, a network partition lets both sides accept writes and believe they are authoritative; when the partition heals there are two conflicting histories and no principled way to reconcile them. Consensus prevents this by refusing to let a minority make progress at all.
And because the cost of that safety is visible in every system built on it. Writes require a round trip to a majority, so latency is set by the slower half of the cluster and, across regions, by geography. The cluster tolerates only (n-1)/2 failures, so five nodes survive two. During an election, there is a window with no leader where writes are refused. These are the trade-offs interviews probe, not the message formats.
How it works
Majority overlap is the safety argument
Any two majorities of the same node set intersect in at least one node. So if a value is committed by one majority and a later majority tries to commit a conflicting value, the shared node already knows about the first — and refuses. This is why quorum means strict majority rather than any fixed count, and why even-sized clusters give no extra fault tolerance.
Terms and leader election
Time is divided into terms, each with at most one leader. A follower that stops hearing heartbeats becomes a candidate, increments the term, and requests votes. A node grants one vote per term and only to a candidate whose log is at least as up to date as its own. Winning requires a majority, so two leaders in one term is impossible. Randomised election timeouts keep candidates from repeatedly splitting the vote.
Log replication and what commit means
Clients send to the leader, which appends and replicates to followers. Once a majority has stored an entry, the leader marks it committed and applies it; followers learn of the commit and apply it too. Committed means durable on a majority — which is exactly why it survives the loss of a minority, including the leader.
Why only an up-to-date candidate can win
Voters refuse candidates whose log is behind theirs. Combined with majority voting, this guarantees any new leader already holds every committed entry — because a committed entry is on a majority, and any winning candidate must have been voted for by a majority, so the sets overlap on a node holding it. That is the safety property that makes leader changes non-destructive.
What it does not solve
Raft assumes crash-stop failures, not malicious ones — a node that lies needs Byzantine fault tolerance, which is a different and far more expensive class of algorithm. Consensus also cannot make progress without a majority, by design: during a partition the minority side correctly refuses to serve rather than risk divergence. And FLP tells us no deterministic algorithm guarantees termination in a fully asynchronous network, which is why real systems rely on timeouts and are live only when the network behaves.
Implementing it
Use odd cluster sizes. Five nodes tolerate two failures; six also tolerate two while adding a node's worth of latency and cost. Even sizes buy nothing.
Size the cluster for the failures you expect, not for throughput. Consensus does not scale writes — every write still goes through the leader and a majority — so adding nodes increases fault tolerance and latency, not capacity.
Keep consensus clusters within a region where possible. Cross-region majorities pay the speed of light on every write, which is frequently the dominant latency in a distributed database.
Treat reads deliberately: a read served by the leader without confirming it is still leader can be stale after a partition. Linearizable reads need a quorum check or a lease, and that cost should be a conscious choice per operation.
// Fault tolerance is (n-1)/2, which is why even sizes are wasteful.
//
// nodes majority failures tolerated
// 3 2 1
// 4 3 1 <- no better than 3, strictly slower
// 5 3 2
// 6 4 2 <- no better than 5
// 7 4 3
//
// Partition of a 5-node cluster into 3 | 2:
// the side of 3 has a majority -> elects a leader, keeps serving
// the side of 2 cannot reach a majority -> refuses writes
// Both sides returning success is what consensus exists to prevent.Interview questions and how to answer them
Why does consensus require a majority rather than any fixed number?
Because any two majorities of the same set must share at least one node. That overlap means a decision committed by one majority cannot be contradicted by another — the shared node would have to accept both and refuses. It is a structural guarantee against split brain, not a redundancy heuristic, and it is why a minority must refuse to make progress.
Walk me through what happens when a Raft leader fails.
Followers stop receiving heartbeats and, after a randomised timeout, one becomes a candidate, increments the term and requests votes. Voters grant at most one vote per term and only to a candidate whose log is at least as current as theirs. A majority wins and begins sending heartbeats. Because a committed entry lives on a majority and the winner was elected by a majority, the sets overlap — so the new leader necessarily holds every committed entry.
A 5-node cluster partitions 3 and 2. What happens?
The side with three has a majority: it elects a leader if needed and continues serving reads and writes. The side with two cannot reach a majority, so it refuses writes and cannot elect a leader. That is the system choosing consistency over availability for the minority, deliberately — both sides accepting writes is precisely the outcome consensus exists to prevent.
Does adding nodes make a consensus cluster faster?
No, slower. Every write still goes through the leader and must reach a majority, and a larger majority means waiting for more acknowledgements, with latency set by the slower members. More nodes buy fault tolerance, and only at odd sizes. Read throughput can be scaled with followers, but only if you accept stale reads or pay for a quorum check.
What kinds of failure does Raft not handle?
Byzantine ones. It assumes nodes crash or become unreachable but do not lie, forge messages or behave arbitrarily. Tolerating that requires BFT algorithms with substantially more nodes and messages. Raft also cannot guarantee termination in a fully asynchronous network — the FLP result — which is why practical implementations depend on timeouts and are live only during reasonable network conditions.
Answers that lose the round
- Saying quorum means "several nodes" rather than a strict majority, and missing the overlap argument
- Using even-numbered clusters, paying for a node that adds no fault tolerance
- Expecting consensus to scale writes — it bounds them to leader plus majority
- Assuming a leader read is linearizable without a quorum check or lease
- Believing Raft tolerates malicious nodes; it assumes crash-stop failures
- Treating the no-leader window during an election as a bug rather than expected behaviour
- Spreading a cluster across regions and being surprised by write latency
FAQ
How is Raft different from Paxos?
They solve the same problem with comparable guarantees. Raft was designed for understandability: a strong leader, terms, and clearly separated subproblems. Multi-Paxos is more flexible and considerably harder to reason about and implement correctly, which is why most modern systems — etcd, Consul, CockroachDB, TiKV — chose Raft.
Are reads from the leader always consistent?
Not automatically. A leader partitioned from the cluster may not yet know it has been replaced, so a naive local read can return stale data. Linearizable reads require confirming leadership with a quorum, or holding a time-based lease. Many systems offer both and make the trade a per-query choice.
What is a learner or non-voting member for?
Replication without participating in quorum. Useful for adding a node — it catches up before voting, so it does not enlarge the majority while behind — and for geographically distant replicas that should serve reads without adding cross-region latency to every write.
Where would I encounter Raft in practice?
etcd, which underpins Kubernetes; Consul; CockroachDB and TiKV per range; Kafka's KRaft controller quorum. Knowing that control-plane metadata typically sits on a consensus cluster explains a lot of operational behaviour — including why etcd latency degrades an entire Kubernetes cluster.