Distributed systems

Stateless vs stateful services

Distributed systemsArchitectureDecision guide

Short answer

Keep request handling stateless by default and place durable state in a purpose-built store. Choose a stateful service when locality, ordered ownership, or a large working set is central to the workload, and then design membership, failover, recovery, and data movement as part of the service contract.

Written and reviewed by Sahil Srivastav

What each one actually is

A stateless instance can answer a request from its inputs and shared stores; any instance can handle the next request. This makes replacement and load balancing straightforward.

A stateful instance owns sessions, partitions, files, streams, or other data that affects future work. Placement and recovery matter because sending a request to another instance may not be equivalent.

State can be durable or merely in-memory. A process-local cache creates state operationally even if losing it is acceptable; model both the loss impact and rebuild time.

Side by side

 Stateless servicesStateful services
Load balancingAny healthy instancePartition ownership or affinity required
RecoveryReplace and replay from shared stateRestore, replicate, or transfer owned state
ScalingAdd interchangeable replicasMove partitions and rebalance data
DeploymentsRolling replacement is simpleDrain, checkpoint, or coordinate ownership
LocalityUsually limitedCan exploit memory, disk, or partition locality
Failure impactRequests fail briefly or retryOwned data or ordering may be unavailable
TestingFocus on shared-store contractsInclude failover, split brain, and recovery tests
Typical examplesHTTP API and render workersDatabases, brokers, stream processors

Choose Stateless services when

  • Requests can be reconstructed from input and shared durable stores
  • Instances should autoscale and be replaced without data movement
  • The service is an API, worker, or gateway with no ownership requirement
  • A shared database or object store already provides the needed durability

Choose Stateful services when

  • Partition ordering or local working state is central to correctness
  • Moving data on every request would cost more than keeping ownership
  • The service is itself the durable store or stream processor
  • The team can operate failover, rebalancing, and recovery procedures

The trade-off in detail

Statelessness moves complexity into shared dependencies. A session store, cache, database, and object store each become part of the request path, so their latency and failure modes still determine availability.

Stateful systems can be faster because data is local, but a restart is a data movement event. Measure recovery time and backlog, and make ownership fencing explicit so an old instance cannot keep writing after failover.

Sticky sessions can hide an accidental stateful design. They improve locality but reduce balancing flexibility and make host loss more visible; use them only when the affinity contract is intentional.

Things that are commonly said and are wrong

  • “Stateless means no data.” It means request correctness does not depend on a particular instance retaining data.
  • “Stateful services cannot scale.” They can shard and rebalance, but scaling includes moving and validating state.
  • “A replicated stateful service is automatically safe.” Replication needs quorum, fencing, recovery, and tested failover semantics.

Decide it in a real repository

Choosing correctly on a whiteboard and enforcing the choice in code are different skills. Gronex ships broken backend repositories whose tests assert the invariant, not the happy path.

FAQ

Should sessions live in the API process?

Usually no. Put sessions in a shared store or use signed, bounded client tokens when appropriate so instances can be replaced and requests can be balanced freely.

Is Redis stateful?

Yes when its data affects correctness or future work, even if it is used as a cache. Decide whether loss is acceptable and configure persistence and recovery accordingly.

How do stateful services deploy safely?

Drain traffic, checkpoint or replicate state, fence the old owner, move ownership if needed, and verify recovery before accepting work. A rolling restart alone is not a state transfer protocol.

Other decisions engineers weigh