PlatformInternal APIsMulti-tenancy

Platform Engineering Interview Questions

Platform engineering interviews have an unusual property: your users are engineers, and they can route around you. That changes what good looks like. An interface that is technically superior and unpleasant to adopt loses to a worse one that is easy, so adoption and migration cost are first-class design concerns rather than afterthoughts.

The second distinctive theme is that your mistakes are multiplied. A bug in a product service affects that service; a bug in a platform primitive affects every team using it, and a breaking change to your API is a coordination problem across the whole organisation. Interviewers probe whether you design with that asymmetry in mind.

Written and reviewed by Sahil Srivastav

What the bar actually is

You should be able to reason about an internal interface as a product: who adopts it, what the migration path is for existing consumers, how you version it, and how you deprecate without stranding anyone. "We would tell teams to migrate" is not a plan, and interviewers will ask what happens to the team that does not.

On multi-tenancy you are expected to understand isolation as something enforced rather than intended. What stops one tenant's load degrading another, what stops one tenant reading another's data, and what the failure mode is when the enforcement is only in application code. The noisy-neighbour problem is the canonical platform question.

How the rounds are structured

Design an internal platform capability

A shared queue abstraction, a deployment primitive, a multi-tenant data layer. Adoption cost and migration path are scored alongside the design.

Multi-tenancy and isolation

Noisy neighbours, per-tenant limits, data isolation, and what happens when one tenant is three orders of magnitude larger than the rest.

Backwards compatibility and deprecation

How you change an interface that fifty services depend on, and how you remove the old one without an organisation-wide freeze.

Implementation round

Usually a library or service with an API surface, where the interface quality is assessed as much as the implementation.

What this interview bar tests

Adoption cost as a design input

The best platform answer includes how a team migrates in an afternoon. Interfaces requiring a quarter of coordinated work do not get adopted, however good they are.

Isolation that is enforced

Per-tenant quotas, partition keys the planner can prune, separate pools. Isolation that exists only as a convention in application code is not isolation.

Deprecation with a real path

Parallel support, usage telemetry to find stragglers, and a date with a consequence. Knowing who still calls the old endpoint is half the problem.

The cost of your own abstraction

Every platform layer adds indirection, a debugging hop and a thing to learn. Being able to say when the abstraction is not worth it is a strong signal.

A preparation plan that works

  1. 1Take an internal interface you have built or used and write its migration story: how a consumer adopts it, how it versions, how the previous thing gets removed. Most candidates have never written this down.
  2. 2Work through the noisy-neighbour problem concretely: how you would detect it, what limit you would impose, and where that limit is enforced. Being specific about the enforcement point is what distinguishes the answer.
  3. 3Practise explaining a platform decision from the consuming team's perspective rather than yours. Platform interviews reward empathy with the adopter, and it is easy to demonstrate and easy to miss.
  4. 4Prepare one example of an abstraction you chose not to build, with the reasoning. It is unusual, it is credible, and it is the clearest evidence of platform judgement.

Questions you should expect

One tenant's traffic is degrading every other tenant. What do you do?

Detect, then isolate, then enforce. Attribute load per tenant first, because the assumption about which tenant is responsible is often wrong. Then impose a per-tenant limit at a point that cannot be bypassed, and check the physical layer — if partitioning does not let the planner prune by tenant, every tenant pays for the largest one regardless of application-level limits.

How do you change an API that fifty services depend on?

Additively and in parallel. Introduce the new shape alongside the old, use telemetry to find every remaining caller, migrate consumers incrementally, then remove the old path with a date and a consequence. A breaking change announced by email is a coordination failure dressed as a plan.

Teams are routing around your platform. Why might that be rational?

Because adoption cost exceeds the benefit for them, the abstraction leaks when they need something it does not expose, or debugging through it is harder than without it. The answer that scores is diagnostic rather than defensive — platform adoption is a signal, and teams going around you is data.

What is the cost of the abstraction you just designed?

An extra hop to debug through, a thing every new engineer must learn, a coupling to your release cycle, and a constraint on consumers who need something you did not anticipate. Platform candidates who can enumerate their own costs are immediately more credible than those who only enumerate benefits.

What gets candidates rejected

  • Designing an interface with no migration path for existing consumers
  • Treating isolation as a convention rather than something enforced at a specific point
  • Announcing breaking changes instead of running parallel support with usage telemetry
  • Being defensive about teams routing around the platform rather than diagnosing why
  • Unable to name any cost of your own abstraction
  • Assuming all tenants are similarly sized, which is almost never true
  • Optimising for platform elegance over consumer ergonomics

What to practise, in order

Hot partition in a multi-tenant database

The noisy-neighbour problem in executable form: correct results, a physical plan that touches every tenant's data, and isolation that has to be enforced rather than intended.

API rate limiting

Per-tenant quota enforcement where the limit must hold under concurrency and across instances.

Feature flag and rollout targeting

A platform primitive with real consumers: targeting rules, consistent bucketing, and change without redeploying.

Zero-downtime database migration

Changing something underneath live consumers without breaking them — the platform problem in miniature.

Practise in a real repository

Gronex ships broken backend repositories with failing test suites that encode the production invariant. You read the evidence, find the defect, and make the tests pass — which is what the round actually measures, rather than whether you can recite a definition.

FAQ

How is platform different from backend in interviews?

The user changes, and so does the scoring. Backend rounds optimise for the end user's correctness and latency; platform rounds optimise for the adopting engineer's ergonomics and migration cost. The same design can score well in one and badly in the other.

Do I need deep Kubernetes knowledge?

Often yes, because it is frequently the substrate. But the probed depth is about what you expose to teams and what you hide, not about API recall — an abstraction that leaks the hard parts and hides the useful ones is the failure mode.

Is multi-tenancy always part of the interview?

In some form, almost always, because it is the defining platform constraint. Even when not named, questions about limits, fairness and isolation are the same question.

How do I show platform impact without owning a platform team?

A shared library others adopted, a build time you cut, a pattern you standardised. The signal is impact through other engineers' work, not the job title that produced it.

Related

Other preparation tracks