Distributed systems

Leader election: interview questions and how to answer them

Leader election gives one node the right to act for a bounded term — and because a leader cannot tell a partition from a slow network, safe designs assume there may briefly be two.

Written and reviewed by Sahil Srivastav

Distributed systemsCoordinationSingleton execution

What it actually is

Leader election is the procedure by which a group of identical nodes agrees that exactly one of them holds a distinguished role: sequencing writes, driving a scheduler, owning a shard. The output is not just an identity but a term — a monotonically increasing number — and the term is what makes the result usable, because it lets everyone else reject instructions from a superseded leader.

Elections are always lease-based in practice. The leader holds authority for a bounded interval and must renew; if it fails to renew, followers start a new election. This gives liveness when a leader dies but means a live leader can lose its lease during a pause or a partition without observing that it has, which is the root of every split-brain incident.

It is worth separating the election from what the leader is for. Electing a leader is a solved problem with well-tested implementations. Making the leader’s side effects safe across a handover — flushing in-flight work, not double-running a job, not writing with stale authority — is the part that is specific to your system and the part interviewers probe.

Why it matters in production

Because “this must run exactly once across the fleet” is an extremely common requirement — nightly reconciliation, a cron-driven billing sweep, a single consumer of an ordered stream — and the naive implementations either run it on every node or run it on none after a deploy. Election turns a fleet of identical replicas into a system with a designated actor without introducing a single point of deployment.

It also matters for latency and correctness in storage systems. A single leader per shard gives a natural serialisation point, which is how systems obtain linearizable writes cheaply: order is decided in one place rather than negotiated per operation. The cost is that the leader is a bottleneck and its failure is a brief write outage.

And the failure mode is expensive. Two nodes that both believe they lead will both advance a scheduler, both accept writes, or both emit notifications — producing duplicate charges or divergent state that must be reconciled by hand.

How it works

Lease plus renewal

The simplest correct-enough scheme: a key in a strongly consistent store, claimed with a compare-and-set and a TTL, renewed at a fraction of the TTL. Followers watch the key and contend when it lapses. etcd leases, ZooKeeper ephemeral sequential nodes, and Kubernetes `Lease` objects are all this pattern, differing mainly in how they detect the holder’s death.

The term number is the safety mechanism

Each successful election increments a term. Every action the leader takes carries its term, and any participant that has seen a higher term rejects it. This is what stops a revived old leader from doing damage: its writes arrive stamped with term 7 into a world that has moved to term 8, and they are refused. Without a term, an election gives you a hint, not a guarantee.

Why a leader cannot know it is still the leader

Authority is granted by others, so knowing you still hold it requires hearing from them — and the absence of a message is indistinguishable from a slow one. A leader can only reason about time since its last successful renewal, using its own clock, which may have been paused or stepped. The practical rule is that a leader must re-verify before any irreversible action, or the action must be idempotent.

Election storms and randomised timeouts

If every follower times out simultaneously, they all campaign, split the vote, and no one wins — repeatedly. Raft solves this by randomising the election timeout per node so one candidate reliably starts first. Systems that hand-roll election with a fixed timeout see exactly this pathology under packet loss: continuous re-elections and no progress.

Handover is the hard part

On losing leadership a node must stop its work promptly: cancel timers, stop consuming, abandon in-flight batches, and close anything that writes. On gaining it, the new leader must assume the previous one may still be finishing something, so it should resume from durable state rather than from memory, and its first actions should be safe to repeat.

Implementing it

Use an existing implementation rather than writing an election: etcd or ZooKeeper if you run them, the Kubernetes lease-based `leaderelection` helper if you are on Kubernetes, or your database’s advisory locks for small deployments. A hand-rolled election on top of an eventually consistent store is an incident waiting for a network blip.

Make the leader’s work idempotent anyway. Election reduces how often two nodes act; idempotency makes it harmless when they do. A nightly job that claims each unit of work with a conditional update is safe under a double election; one that assumes exclusivity is not.

Stop work on demotion with the same care as starting it on promotion — a callback that logs “lost leadership” but leaves the scheduler thread running is a very common bug, because it looks correct in tests where demotion never happens.

Alarm on leadership churn. Frequent handovers usually mean the renewal interval is too close to the lease TTL for the observed network latency or the node is GC-pausing past it, and both degrade throughput long before anyone notices a correctness problem.

Interview questions and how to answer them

How would you make a scheduled job run on exactly one node in a fleet of ten?

Elect a leader via a lease in a strongly consistent store and have only the leader schedule. Then assume that is imperfect: each unit of work is claimed with a conditional update — `UPDATE tasks SET owner = $1, state = 'RUNNING' WHERE id = $2 AND state = 'PENDING'` — so a second leader during a handover window finds nothing to claim. Election gives efficiency, the claim gives correctness.

Why does a leader need a term number?

Because leadership is revocable and a revoked leader may not know it. The term lets every other participant order authority: anything stamped with an older term is rejected outright. Without it, a leader that was partitioned for thirty seconds comes back and issues commands that look exactly as valid as the current leader’s.

Can two nodes believe they are leader at the same time?

Yes, briefly, in any lease-based system — the old leader has not yet noticed its lease lapsed while the new one has already won. A correct design does not try to prevent that window; it makes the window harmless, by having the storage layer reject writes from an older term and by making the leader’s actions idempotent.

What are randomised election timeouts for?

To break symmetry. With identical timeouts, all followers campaign at the same instant, split the vote, and repeat, so the cluster can stay leaderless indefinitely under mild packet loss. Randomising over a range makes one node campaign first with high probability, so an election terminates in one or two rounds.

Your service logs a leadership change every few minutes. Is that a problem?

Yes, even if nothing has visibly broken. Each handover is a brief write stall plus a window where two nodes may act, so frequent churn multiplies the exposure. The usual causes are a renewal interval too close to the lease TTL for the observed latency, GC pauses on the leader exceeding the renewal deadline, or the renewal running on a thread pool saturated by application work.

How does leader election relate to consensus?

Election is one of the things consensus provides, and consensus is how election is made safe. Raft elects a leader per term using majority votes, and the same majority requirement guarantees at most one leader per term. You can also elect a leader *using* an existing consensus system, which is what a lease in etcd or ZooKeeper is — you are borrowing their Raft or ZAB rather than running your own.

Answers that lose the round

  • Describing election without terms or fencing, so a revived old leader is indistinguishable from the current one
  • Electing with a lock in an asynchronously replicated store, where a failover can hand the same lock to two nodes
  • Using a fixed election timeout on every node, causing split votes and repeated elections under packet loss
  • Assuming the leader knows when it has lost the lease — it can only know when it last renewed successfully
  • Treating election as sufficient for once-only execution, instead of also making the work idempotent
  • Renewing on a timer that shares a thread with the leader’s workload, so a busy leader loses its lease spuriously
  • Forgetting to stop in-flight work on demotion, which is precisely how two active leaders produce duplicate effects

Practise in a real repository

Explaining a concept and enforcing it in code are different skills, and machine coding rounds test the second. Gronex ships broken backend repositories whose test suites assert the invariant rather than the happy path.

FAQ

Is a Kubernetes Deployment with one replica the same as leader election?

No. During a rolling update or a node drain, the old pod can still be running while the new one starts, and a partitioned kubelet can leave a pod alive that the control plane considers gone. `replicas: 1` is a scheduling intent, not a mutual-exclusion guarantee, which is why controllers that need singleton behaviour use a `Lease` object even at one replica.

Can I elect a leader using a database row?

Yes, and for modest systems it is a good choice: a single-row lease with a conditional update, or a Postgres advisory lock held on the leader’s connection. The connection-scoped advisory lock has a nice property — it disappears when the connection dies — but it binds leadership to a connection, so a network hiccup causes a handover.

Should followers do anything useful?

Usually yes, and designing that is worthwhile. Followers can serve reads with a stated staleness bound, keep warm caches, and pre-validate work so a promotion is fast. What they must not do is anything the leader is also doing, which is why the split between leader-only and any-node work needs to be explicit in the code rather than implied.

Related

More backend concepts