Performance

Horizontal vs vertical scaling

PerformanceInfrastructureDecision guide

Short answer

Scale vertically first when one machine has headroom and simplicity matters. Scale horizontally when a single host limits capacity, availability, or deployment safety, and make state, load balancing, coordination, and data partitioning explicit before adding replicas.

Written and reviewed by Sahil Srivastav

What each one actually is

Vertical scaling increases the CPU, memory, storage, or network capacity of one machine. It preserves local calls and usually needs less application change, but every host has a finite ceiling and upgrades can be disruptive.

Horizontal scaling adds machines and distributes requests or data across them. It improves capacity and failure tolerance when the workload is partitionable, but introduces routing, replication, consistency, and operational coordination.

Scaling the application tier and scaling the database are separate decisions. Stateless web replicas are easy to add; a write-heavy database may need a different design entirely.

Side by side

 Horizontal scalingVertical scaling
Application changeOften smallRequires partitionable work and shared state design
Failure toleranceOne larger failure domainCan survive host loss with redundancy
Capacity ceilingBound by largest machineAdds nodes while the workload can partition
CoordinationMostly localLoad balancing, membership, replication, or consensus
State handlingLocal state is straightforwardSessions and files need shared or routed storage
LatencyLocal calls stay cheapNetwork hops and tail latency increase
Cost shapeSimple but expensive at high tiersEfficient at scale, with platform overhead
Upgrade pathResize or replace a hostRolling changes are possible if replicas are safe

Choose Horizontal scaling when

  • The workload is still below a single host’s practical capacity
  • The service has local state or an unpartitioned database
  • The team needs the fastest low-risk capacity increase
  • A larger machine remains cheaper than operating a distributed fleet

Choose Vertical scaling when

  • Traffic or data exceeds one host’s safe limit
  • Requests can be served independently by replicas
  • Availability requires surviving a host or zone failure
  • Rolling capacity and upgrades matter more than local simplicity

The trade-off in detail

Horizontal replicas do not make a stateful service stateless. Sessions, uploads, caches, scheduled jobs, and connection pools must be shared, partitioned, or deliberately sticky. A load balancer only spreads requests; it does not solve ownership.

Vertical upgrades can postpone a redesign and are often the right operational decision. They also create a cliff: once the largest host is full, moving to horizontal architecture under load is harder than planning the boundary earlier.

A database may scale reads horizontally while writes remain central. Measure the saturated resource and the invariant that prevents partitioning before promising linear scale.

Things that are commonly said and are wrong

  • “Horizontal scaling is always cheaper.” Fleet, network, storage, and operator costs can outweigh a larger host at modest scale.
  • “More replicas automatically improve performance.” Lock contention, a shared database, or a hot partition can remain the bottleneck.
  • “Vertical scaling means downtime.” Many providers support online resizing, replicas, or failover, though the exact availability contract must be tested.

Decide it in a real repository

Choosing correctly on a whiteboard and enforcing the choice in code are different skills. Gronex ships broken backend repositories whose tests assert the invariant, not the happy path.

FAQ

Which should a startup choose?

Use vertical scaling while one host meets the measured workload and availability target. Add horizontal replicas when traffic, failure tolerance, or deployment needs justify the added state and coordination work.

Can a database scale horizontally?

Reads commonly scale with replicas. Writes require partitioning, sharding, distributed transactions, or a different data model, each with consistency and operational costs.

How do I know it is time to scale out?

Look for sustained saturation, tail latency, failure-domain requirements, or a capacity ceiling that vertical upgrades cannot safely address. Do not scale out from request count alone.

Other decisions engineers weigh