Linux / shell
Readiness probe failing while the application looks healthy
Written and reviewed by Sahil Srivastav
Readiness probe failed: HTTP probe failed with statuscode: 503What this error actually means
Readiness answers whether this pod should receive ordinary service traffic now. A failing readiness probe marks the container unready; it does not itself restart it. With normal Service endpoint behaviour, unready endpoints are excluded from traffic, so a process can be Running and accept a local request while receiving no requests through its Service.
A successful manual request is useful only if it tests the same contract. Kubelet’s HTTP probe normally targets the pod IP and configured port, not the ingress URL. An ingress can rewrite paths, terminate TLS and add authentication headers. Port-forwarding can also succeed against a process listening only on loopback even when node-to-pod traffic cannot reach it.
A 503 from the probe proves that an HTTP server responded; it points towards the endpoint’s readiness decision. Connection refused, timeout, 401 and 404 mean different things. Treating every failure as a timeout problem removes this distinction and can delay a rollout without improving service availability.
Causes, most common first
- 1Probe and application disagree on address, port or path. The server binds to 127.0.0.1, the probe uses the service port instead of the container’s listening port, or a context-path change moved the endpoint. An ingress route working says little about a direct pod-IP request.
- 2Authentication or virtual-host routing intercepts the check. A new middleware layer returns 401, or the health handler exists only for a specific Host header. Do not embed a privileged user token in the probe; define a minimal health contract with intentional access controls.
- 3Readiness depends on too much shared infrastructure. The endpoint synchronously checks every downstream service on every probe. One optional dependency fails and every replica removes itself from service, even though most operations could still succeed. The resulting zero-capacity state can amplify the original outage.
- 4The health request competes with saturated application work. A full request queue, blocked event loop or CPU throttling delays an otherwise cheap endpoint. If ordinary requests are equally delayed, the pod is not actually healthy; merely increasing the probe timeout disguises overload rather than restoring capacity.
When you see it
- The pod stays Running but shows fewer ready containers than expected
- Requests through port-forward succeed while the Service has no ready endpoints
- Readiness flaps during CPU throttling or a slow shared dependency
- Application logs look normal because the health endpoint rejects before normal handlers
How to diagnose it
Step 1
Read the exact probe result and live configuration
Replace demo and app-pod with the affected namespace and pod. Identify which container failed and whether the result was an HTTP status or a connection failure. Inspect the admitted probe fields rather than assuming the chart defaults are still in effect.
kubectl -n demo describe pod app-pod
kubectl -n demo get pod app-pod -o jsonpath='{range .spec.containers[*]}{.name}{"\t"}{.readinessProbe}{"\n"}{end}'Step 2
Reproduce the direct request without ingress behaviour
From an existing approved diagnostic environment with network access, call the pod IP with the configured scheme, port, path and Host header. The example IP and path are illustrative. A debug pod is not the kubelet’s network origin, so a policy-related discrepancy still needs node-side investigation.
curl -i --max-time 2 http://10.42.1.15:8080/readyzStep 3
Check whether the endpoint was removed as expected
Inspect EndpointSlices for the actual Service, here app. Compare addresses and readiness conditions with the affected pod. Custom routing and publishNotReadyAddresses can alter the normal behaviour; absence of traffic is not proof of a selector bug.
kubectl -n demo get endpointslice -l kubernetes.io/service-name=app -o yamlStep 4
Measure health-handler work during a failure
Record the endpoint’s decision reason and elapsed time without returning secrets to callers. Correlate failures with CPU throttling, queue depth and dependency latency. A readiness handler waiting for a database connection from an exhausted pool has inherited the same bottleneck it is meant to describe.
The fix
Align the probe with an explicit application endpoint on the actual listening port. Bind the service to its intended pod-facing interface, and use named container ports consistently. If host-based routing is required, configure the appropriate HTTP Host header rather than sending the probe to a different external hostname.
Make readiness express ability to serve the workload’s required operations. Keep optional dependencies out of an all-or-nothing gate when a supported degraded mode exists. A required datastore may legitimately make the application unready, but choose that behaviour consciously and test the all-replicas-affected case.
Bound the work done by the health check. Prefer a cheap local assessment of initialisation and recent dependency state over unbounded synchronous fan-out. Ensure the state expires appropriately; a permanently cached healthy result is no longer a useful readiness signal.
Resolve real overload through admission control, concurrency bounds and sufficient capacity. Adjust timeout and failure thresholds from measured healthy behaviour to absorb brief jitter, then verify that a truly unavailable pod is still removed within an acceptable window.
How to stop it coming back
- Test probes against the built container without ingress rewrites or developer credentials
- Log a bounded, non-sensitive reason for transitions to unready
- Simulate loss of one shared dependency and observe available service capacity
- Monitor ready replica count alongside restart count; readiness failures may never restart
FAQ
Does a failed readiness probe restart the pod?
No. Readiness controls availability for routing. Liveness or startup probe failure can trigger a container restart after its failure threshold, which is why using one expensive endpoint for all probe types is risky.
Why does localhost return 200?
It may reach a loopback-only listener or bypass host, network and routing conditions used by kubelet. Compare the exact requests before concluding Kubernetes is incorrectly marking the pod unready.
Should the check always return 200 to keep traffic flowing?
No. That sends work to instances unable to complete it. Define the required serving capabilities precisely, support deliberate degraded operation, and make readiness reflect that contract.
Related
Other errors engineers hit next to this one
- java.util.ConcurrentModificationException
- OutOfMemoryError: unable to create new native thread
- RejectedExecutionException: Task rejected from ThreadPoolExecutor
- Found one Java-level deadlock (thread dump)
- FATAL: sorry, too many clients already
- Sessions stuck in "idle in transaction"
- ERROR: deadlock detected
- ERROR: canceling statement due to statement timeout