Linux / shell

Liveness probe killing a slow-starting application

Written and reviewed by Sahil Srivastav

KubernetesStartup probesCold starts
Liveness probe failed: HTTP probe failed with statuscode: 503

What this error actually means

A liveness probe asks whether restarting this container can recover a broken process. When it runs before legitimate initialisation finishes, it mistakes not-yet-started for no-longer-working. Kubelet terminates the process, initialisation begins again, and the application never reaches the state that would satisfy the check.

Readiness does not protect a process from liveness. They are independent checks unless startup probing gates them. A startup probe gives each container start a finite initialisation window; after it succeeds, the normal liveness and readiness checks take over. This separates a potentially long cold start from the shorter recovery time desired for a later deadlock.

The same apparent loop can also come from a genuine startup hang or an incorrectly addressed endpoint. Extending the budget is justified when logs show forward progress towards a known completion point. It is not a cure for waiting forever on a schema lock, a broken DNS name or a health URL that will always return 404.

Causes, most common first

  1. 1A steady-state liveness deadline was applied to cold start. Class loading, dependency injection, model loading or cache reconstruction takes longer than the allowed initial failures. The probe configuration may have been copied from a much smaller service rather than measured against this workload.
  2. 2CPU throttling makes startup much slower than development. Initialisation can briefly need substantial CPU even when steady-state usage is small. A low CPU limit stretches that burst over a longer wall-clock interval. Comparing average CPU after startup misses the exact interval in which the probe kills the process.
  3. 3Liveness includes a remote dependency. A database outage or slow external call makes the health endpoint fail even though the process could recover when the dependency returns. Restarting every replica adds connection storms and discards useful local state without repairing the remote service.
  4. 4Initialisation never converges. A migration lock, unbounded retry or deadlock prevents completion. Repeatedly enlarging the window delays detection while preserving the failure. Establish a progress marker and a finite dependency deadline before deciding that startup is merely slow.

When you see it

  • Each attempt reaches roughly the same startup stage before a Killing event
  • The service starts locally but fails under the cluster’s CPU or memory limits
  • Warm restarts succeed while first starts after a rollout or cache miss fail
  • A temporary generous startup window allows initialisation to finish and stay healthy

How to diagnose it

Step 1

Prove that the probe initiated termination

For the actual namespace and pod, compare Unhealthy and Killing events with previous-container logs. If the process exited before the probe failures, investigate the application exception instead. A lastState exit alone cannot establish who initiated shutdown.

kubectl -n demo describe pod app-pod
kubectl -n demo logs app-pod -c app --previous --timestamps

Step 2

Read all three probe definitions together

Identify initial delays, periods, timeouts and failure thresholds. Check whether a startup probe is present and whether it succeeds too early, for example when a management port opens before the application can make progress.

kubectl -n demo get pod app-pod -o jsonpath='{range .spec.containers[*]}{.name}{"\n"}{.startupProbe}{"\n"}{.livenessProbe}{"\n"}{.readinessProbe}{"\n"}{end}'

Step 3

Measure cold-start completion under the deployment limits

In staging, use the same image, resources and representative data with a finite but generous startup window. Record timestamps for process entry, dependency initialisation and readiness. Include cold caches and node contention rather than using the fastest observed attempt as the budget.

Step 4

Separate useful work from waiting

Inspect runtime profiles or thread stacks while startup is slow. CPU-bound class loading calls for a different change from a thread waiting indefinitely for a migration lock. Correlate CPU throttling metrics with the startup interval instead of inferring throttling from low average utilisation.

The fix

Add a startup probe whose success means the process has completed the initialisation needed for ordinary health assessment. Choose failureThreshold times periodSeconds as an approximate startup budget with measured headroom; probe duration and scheduling mean this is not a precise application deadline.

Keep steady-state liveness narrow: detect a condition for which restarting this process is useful. Readiness handles temporary inability to serve. If a shared dependency disappears, repeatedly restarting healthy local processes generally makes recovery slower.

Reduce avoidable startup work and give essential initialisation realistic CPU resources. Move coordinated schema changes out of every replica’s startup path where possible. If preparation must happen first, a bounded init-container task makes that dependency explicit and separately observable.

Validate both sides of the policy: a cold process must survive a legitimate start, while a deliberately wedged process must still restart within the intended recovery interval. A configuration that makes both wait ten minutes has hidden one failure by weakening another guarantee.

# Container spec fragment; implement these endpoints first.
ports:
  - name: http
    containerPort: 8080
startupProbe:
  httpGet:
    path: /startupz
    port: http
  periodSeconds: 5
  timeoutSeconds: 2
  failureThreshold: 36
livenessProbe:
  httpGet:
    path: /livez
    port: http
  periodSeconds: 10
  timeoutSeconds: 2
  failureThreshold: 3
readinessProbe:
  httpGet:
    path: /readyz
    port: http
  periodSeconds: 5
  timeoutSeconds: 2
  failureThreshold: 3

How to stop it coming back

  • Track startup duration separately for cold nodes, warm nodes and data-size changes
  • Run rollout tests with production CPU limits and realistic initialisation data
  • Require a reason why restarting helps every condition included in liveness
  • Record named startup stages so a stalled stage is distinguishable from slow progress

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Why not just increase initialDelaySeconds?

A fixed delay can work for simple predictable startup, but it postpones checking even when startup finishes early. A startup probe observes completion and then permits the normal checks to operate, while retaining a finite failure budget.

Does a startup probe replace readiness?

No. Startup success happens once per container start. Readiness continues to assess whether the running application can accept traffic as dependencies and local capacity change.

Is a longer timeout the fix for every liveness failure?

No. A wrong path remains wrong, a deadlock remains stuck, and a remote outage remains remote. Increase timing budgets only after measuring legitimate work that the old settings interrupted.

Related

Other errors engineers hit next to this one

Full error and symptom index →