Distributed systems
Service discovery: interview questions and how to answer them
Service discovery answers “where is an instance of X right now”, and its hard part is not registration but timely, safe removal of instances that have stopped being healthy.
Written and reviewed by Sahil Srivastav
What it actually is
Service discovery replaces hard-coded addresses with a lookup. Instances register themselves — or are registered by the platform — into a registry keyed by logical service name, and callers resolve that name to a current set of endpoints. The registry is thus a piece of shared mutable state in the request path of everything, which is why its consistency and availability properties matter so much.
There are two shapes. In client-side discovery the caller fetches the endpoint list and picks one itself, so it can do locality-aware or load-aware selection with no extra hop. In server-side discovery the caller talks to a stable address — a load balancer, a mesh sidecar, a Kubernetes Service — and that component resolves and balances. Client-side gives control and couples every client to the registry; server-side centralises the logic and adds a hop.
Registration is easy; deregistration is where systems fail. An instance that crashes cannot deregister itself, so the registry must infer death from missing heartbeats or failing health checks, and every layer of caching between the registry and the caller extends the window during which traffic is still sent to a dead address.
Why it matters in production
Because in an autoscaled environment endpoints change constantly. Deploys replace every instance, scale-downs remove them, spot interruptions kill them without notice. Any staleness in the endpoint set turns directly into connection errors or, worse, timeouts — and timeouts consume a caller’s concurrency budget in a way immediate refusals do not.
The registry’s availability is also a systemic risk. If callers cannot resolve a name when the registry is down, a registry outage becomes a total outage. Designs that survive this cache the last known good set and keep using it, treating the registry as a source of updates rather than a dependency for every call — a deliberately availability-over-freshness choice.
And discovery is where graceful shutdown either exists or does not. The correct sequence — deregister, wait for propagation, drain in-flight requests, then exit — is a handful of lines that eliminate the error spike every deploy otherwise produces.
How it works
Registry with TTL heartbeats
An instance registers with a TTL and refreshes periodically; missing refreshes expire the entry. This detects hard crashes without any active probing, and the detection delay is bounded by the TTL. The trade is chattiness proportional to fleet size, and a tuning trap: a TTL short enough for fast detection causes spurious expiry whenever the instance or the registry has a latency hiccup.
Active health checks
The registry or balancer probes each instance. This catches the harder cases a heartbeat misses — a process alive but unable to serve, a saturated thread pool, a broken downstream dependency. Two probes are worth distinguishing: liveness (restart me) and readiness (stop sending me traffic). Conflating them gets healthy-but-busy instances killed instead of merely drained.
Why DNS-based discovery lags
DNS answers carry a TTL, but the caching layers do not all respect it: resolvers round it up, some JVM configurations cache positive lookups for the process lifetime unless `networkaddress.cache.ttl` is set, and connection pools hold sockets to already-resolved addresses indefinitely. So a removed endpoint can keep receiving traffic long after the record changed. If you use DNS, you must also bound connection lifetime.
Graceful shutdown ordering
On SIGTERM: mark readiness failing or deregister, then keep serving for at least one propagation interval, then stop accepting new connections, drain in-flight requests to completion, and exit. Exiting immediately on SIGTERM is the single most common cause of deploy-time 502s, because the endpoint set has not yet propagated to callers.
Consistency model of the registry
Eureka deliberately chooses availability: it serves possibly stale membership and expects clients to retry past bad endpoints. Consul and etcd are consensus-backed and will refuse writes on a minority partition, giving a more accurate view at the cost of availability. Neither is wrong; the choice determines whether your failure mode is “traffic to a dead instance” or “cannot register during a partition”.
Implementing it
Combine a short propagation path with client-side resilience: retry another endpoint on a connection failure, mark failing endpoints as unhealthy locally with an ejection window, and re-resolve periodically rather than once at startup.
Cache the last known good endpoint set on the client and keep serving from it if the registry is unreachable. Fail-open on discovery is almost always correct — an out-of-date list is more useful than no list.
Separate readiness from liveness and give readiness a dependency-aware implementation that does not cascade: if readiness fails whenever a downstream is slow, a downstream blip removes your entire fleet from rotation at once.
Bound connection lifetime (max connection age or max requests per connection) wherever discovery flows through DNS, so pooled sockets to retired endpoints are recycled rather than held for hours.
Interview questions and how to answer them
Client-side or server-side discovery — how do you choose?
Client-side when you want locality-aware or load-aware selection without an extra hop, and you can accept every client embedding a discovery library and a consistent policy. Server-side when you want one place to change routing and balancing policy and to keep clients dumb, accepting the added hop and the balancer as a dependency. A service mesh is the hybrid: server-side logic in a sidecar, so the hop is local.
An instance is terminated. Trace what has to happen before traffic stops reaching it.
The registry must notice — via deregistration if the shutdown was graceful, otherwise via heartbeat expiry or failed health checks. The updated set must propagate to callers or balancers, bounded by their poll or watch latency. Callers must stop using cached resolutions and, critically, must close existing pooled connections to that address. If any of those steps is unbounded — a DNS cache with no expiry, a connection pool with no max age — traffic keeps flowing to a dead endpoint indefinitely.
Why does a deploy cause a burst of 502s, and how do you stop it?
Because pods or instances stop accepting connections before the load balancer knows they are gone. Fix the ordering: on SIGTERM, fail readiness first, then sleep for at least one readiness-probe interval plus the balancer’s propagation delay, then stop the listener and drain. On Kubernetes that means a `preStop` sleep and a `terminationGracePeriodSeconds` longer than the drain time.
Your service registry goes down. What happens to traffic?
It should keep flowing. Clients or sidecars serve from their last known endpoint set and continue retrying the registry in the background; new instances cannot join and dead ones cannot be removed, so the set drifts, but existing capacity keeps serving. A design where resolution failure fails the request converts a control-plane outage into a data-plane outage, which is the mistake worth naming.
How should readiness checks be implemented?
Cheap, local, and dependency-light. Confirm the process can accept work — listener bound, thread pool not exhausted, caches warmed if warm-up is required — and avoid failing on downstream health. If a critical dependency is down, failing readiness fleet-wide removes all capacity and prevents recovery; better to serve degraded responses and let a circuit breaker handle the dependency.
How does Kubernetes do service discovery?
The control plane watches pod readiness and maintains EndpointSlices for each Service; kube-proxy or a CNI programs the node’s dataplane so the Service’s cluster IP load-balances to ready pods, and CoreDNS resolves the Service name. Discovery is therefore server-side and push-driven via watches, which is why propagation is typically fast — but the programming of iptables or IPVS rules and any external load balancer still lag by a small, real amount, which is why the `preStop` drain still matters.
Answers that lose the round
- Resolving a hostname once at startup and holding the address for the process lifetime
- Exiting immediately on SIGTERM instead of deregistering, waiting for propagation, and draining
- Treating the registry as a hard dependency, so its outage becomes a full request-path outage
- Using one health endpoint for both liveness and readiness, so busy instances get restarted
- Making readiness depend on downstream health, which removes the whole fleet from rotation when a dependency is slow
- Setting a very short DNS TTL and assuming clients honour it — many resolvers and runtimes do not
- Relying on heartbeat expiry alone, which cannot detect a process that is alive but unable to serve
FAQ
Is a load balancer service discovery?
It is the server-side form of it. The balancer holds the membership and the callers hold one stable address. The discovery problem does not disappear — something still has to add and remove targets based on health — it simply moves into the balancer and its target-group registration.
Do I need Consul or Eureka if I run on Kubernetes?
Usually not for in-cluster discovery, which Services and EndpointSlices already cover. A separate registry earns its place when you must discover across clusters, across clouds, or include workloads that are not on Kubernetes at all — and in that case it is the multi-environment membership, not the discovery mechanism, that you are buying.
How short should a DNS TTL be for service discovery?
Short (a few seconds to a minute) but never trusted. Treat it as a hint, and pair it with periodic re-resolution inside the client, connection max-age so sockets to retired addresses are recycled, and retry-on-another-endpoint behaviour. The TTL bounds the resolver cache, not the caller’s connection pool.