SDE3SeniorDesign judgement

SDE3 Backend Interview Preparation

By SDE3 the implementation question has been settled — nobody doubts you can write the service. What is being assessed is judgement: which problem you chose to solve, what you deliberately left out, and whether you can defend both when pushed by someone who knows the trade-off space.

The practical shift is that there is no longer a single right answer being graded. An SDE3 round is closer to a technical argument than an exam. You propose, the interviewer applies pressure — "what happens at ten times the traffic", "what if that service is down for an hour", "why not just use a queue" — and the signal is the quality of your reasoning under that pressure.

Written and reviewed by Sahil Srivastav

What the bar actually is

The bar is owning an ambiguous problem end to end. You should be able to take a vague requirement, narrow it into something buildable, state the assumptions you made and why, and identify the two or three decisions that actually matter versus the dozen that do not. Candidates who treat every decision as equally weighty read as unable to prioritise.

You are also expected to think about failure as a first-class concern rather than an afterthought. At SDE2 it is enough to handle the failure cases you are asked about; at SDE3 you are expected to raise them unprompted — what happens on a partial write, a duplicate delivery, a slow dependency, a deploy mid-transaction — and to have a defensible position on which of those you would actually engineer for.

How the rounds are structured

A harder machine coding round with explicit extension

The problem is larger, the time is similar, and the extension is announced up front: build this so that a second payment provider can be added. Your structure is the deliverable as much as your working code.

A system design round with real pressure

Not a recitation of boxes and arrows. Expect to be pushed on one specific part of your design until you either defend it with a reason or concede. Conceding well — "you are right, that fails under X, here is what I would change" — scores better than defending badly.

A debugging or incident round

Increasingly common at this level: here is a production symptom and some evidence, tell us what you would look at. It is the round that most reliably separates people who have operated systems from people who have only built them.

Technical depth on your own past work

A long conversation about something you actually shipped, probing for what went wrong, what you would do differently, and whether you understood the system or just your part of it.

What this interview bar tests

Choosing what not to build

The strongest SDE3 signal is a candidate who narrows scope deliberately and says why: "I am not handling multi-currency because nothing in the requirement implies it, and adding it now would couple the ledger to a rate source."

Failure modes raised unprompted

Partial writes, duplicate deliveries, timeouts without retries, a deploy landing mid-transaction. Naming these before being asked is the difference between senior and mid-level in most rubrics.

Extensibility at the right seam

Abstraction is not free. The signal is putting one seam where change is actually likely, and declining to abstract the four places where it is not — then being able to explain the asymmetry.

Operational consequences of design

How you would deploy it, what you would monitor, what the first alert would be, how you would roll it back. Design that cannot be operated is not finished design.

A preparation plan that works

  1. 1Take three systems you have built and write, for each, the two decisions that mattered and the alternatives you rejected. Being fluent about your own trade-offs is most of the depth round, and most candidates have never articulated them out loud.
  2. 2Practise being pushed. Have someone — or a model — attack one component of your design repeatedly until you either justify it concretely or change it. The skill being built is staying analytical rather than defensive.
  3. 3Work debugging problems specifically. Given a symptom and evidence, narrow to a cause. This is a distinct skill from building, it is increasingly tested at this level, and it is the one most senior candidates have never practised deliberately.
  4. 4Build one problem twice: once naively, once with a seam for the extension you were told was coming. Comparing the two teaches where abstraction pays and where it is dead weight.

Questions you should expect

You need to write to your database and publish an event atomically. How?

You cannot, directly — there is no transaction spanning both. The defensible answer is the transactional outbox: write the intent to a table in the same transaction as the business data, then let a separate worker publish and mark it sent. Then volunteer the consequence: delivery becomes at-least-once, so consumers must be idempotent. Claiming you would use a distributed transaction, or that your broker gives exactly-once, is the answer that ends the round badly.

This endpoint is slow at p99 but fine at p50. Where do you look?

A tail problem is usually a queueing or contention problem rather than a slow-code problem. Look for a bounded resource — a connection pool, a thread pool, a lock on a hot row — and for work whose cost scales with data volume for some users only. Averages hide all of this, which is why the p50 looks healthy.

Why would you not add a cache here?

Because a cache adds an invalidation problem, a staleness window, and a second failure mode in exchange for latency you may not need. Name the condition that would change your mind — read-heavy access on data that tolerates being seconds stale — and what you would measure before deciding.

How do you make this deployable without downtime?

Decouple schema change from code change, and make each step independently safe: additive migration first, code that tolerates both shapes, backfill, then narrow. Mention the lock behaviour — DDL waiting on a long transaction queues every reader behind it, so migrations need a short lock timeout and a retry rather than a long wait.

What gets candidates rejected

  • Designing for scale that was never stated, instead of asking what the actual load is
  • Abstracting every component equally, which signals an inability to judge where change is likely
  • Treating failure handling as a section to mention rather than a property of the design
  • Defending a flawed decision after the interviewer has shown it breaks, rather than conceding and adapting
  • Describing past work only at the level of your own tickets, with no view of the surrounding system
  • Claiming exactly-once delivery, or a distributed transaction across a database and a broker, as though they were available
  • Reciting a design template rather than reasoning about this specific problem

What to practise, in order

Transactional outbox implementation

The dual-write problem, as a real repository: atomic enqueue, at-least-once publish, idempotent consumption, and crash recovery mid-publish. The question above, made executable.

Zero-downtime database migration

Ship a schema change while reads and writes continue. The tests run traffic throughout and assert no failed request and no lost write, so a migration that merely completes does not pass.

Payment ledger consistency

Money that must balance under concurrency, retries and partial failure. The problem where "mostly correct" is obviously not correct.

Read replica consistency failure

Writes to the primary, reads from a lagging standby, and a broken read-your-own-writes guarantee. Practice for the staleness trade-off you will be asked to defend.

Practise in a real repository

Gronex ships broken backend repositories with failing test suites that encode the production invariant. You read the evidence, find the defect, and make the tests pass — which is what the round actually measures, rather than whether you can recite a definition.

FAQ

How is the SDE3 machine coding round different from SDE2?

The problem is larger and the extension is explicit rather than implied. At SDE2 you are asked to build it well; at SDE3 you are told a second variant is coming and judged on whether your structure absorbs it without a rewrite. Working code remains the entry ticket at both levels.

Do I need to know distributed systems theory in depth?

You need the consequences, not the proofs. Know what at-least-once delivery forces on a consumer, why read-your-own-writes breaks on a replica, what a quorum buys and costs. Being able to state CAP precisely matters much less than knowing what you would actually do when a dependency is unavailable.

What if I disagree with the interviewer?

Say so, with a reason, and stay open. A well-argued disagreement is a strong signal; most interviewers are probing to see whether you fold or dig in irrationally. What does not work is defending a position after it has been concretely shown to fail.

How much does the debugging round matter?

More every year, and it is the round senior candidates most often walk into cold. It is hard to fake — you either have the habit of reasoning from evidence to a narrow hypothesis, or you guess and change things. Practising it deliberately is unusually high-leverage.

Related

Other preparation tracks