Linux / shell
CrashLoopBackOff: diagnose the exit behind the backoff
Written and reviewed by Sahil Srivastav
CrashLoopBackOffWhat this error actually means
CrashLoopBackOff describes a waiting interval between repeated container restarts. It is a kubelet behaviour, not the exception that broke the application. The important evidence sits one level earlier: the terminated process’s exit status, its previous logs, or an event showing that a probe asked kubelet to kill it.
The word crash is misleading when the command exits successfully. A Deployment normally expects a long-running process and uses restartPolicy Always. A migration command that completes with exit code zero, or a server that daemonises and lets its original process exit, can be restarted repeatedly even though it never threw an exception.
Backoff protects the node and dependencies from an uncontrolled restart storm. Its timing can vary with Kubernetes configuration and version; do not infer a root cause from the precise delay. Deleting the pod may reset the visible cycle while throwing away the record you needed to diagnose it.
Causes, most common first
- 1Application startup fails deterministically. A malformed configuration, unsupported runtime option or unavailable required dependency makes every attempt exit at the same point. Compare the first failing revision with the last working image and configuration; changing restart timing cannot make invalid input valid.
- 2Memory enforcement or a liveness probe terminates the process. These failures are outside the application’s normal exception path. OOMKilled implicates memory; repeated Unhealthy and Killing events implicate the probe. A shutdown stack trace can be a consequence of the kill rather than the initiating problem.
- 3The container command has the wrong lifecycle. A CLI finishes, a shell backgrounds the server, or the executable starts in daemon mode. Kubernetes tracks the container’s main process, not the existence of an unrelated child after that process exits.
- 4A dependency outage is coupled to process survival. An application exits on each failed database connection, turning a dependency incident into repeated cold starts. Bounded reconnect logic may be appropriate after startup; pretending the dependency is healthy is not. Separate serving readiness from recoverable connection loss.
When you see it
- The restart count rises while the container alternates between Running and Waiting
- Current logs contain only startup lines, but --previous shows the fatal event
- A rollout stalls because replacement pods never remain ready long enough
- The application starts normally and is killed at a repeatable probe deadline
How to diagnose it
Step 1
Inspect the affected container’s last termination
Use the actual namespace, pod and container in place of demo, app-pod and app. In a multi-container pod, the main application can be healthy while a sidecar loops. Init-container status is separate and should be inspected when the application never starts.
kubectl -n demo get pod app-pod -o jsonpath='{range .status.containerStatuses[*]}{.name}{"\t"}{.restartCount}{"\t"}{.lastState.terminated}{"\n"}{end}'
kubectl -n demo describe pod app-podStep 2
Read logs from the instance that exited
The new process may not have reached the failing code yet. Previous logs normally cover only the previous terminated instance, so export them promptly and use central logging for older attempts. Empty logs also matter: execution or configuration may fail before application logging initialises.
kubectl -n demo logs app-pod -c app --previous --timestamps --tail=200
kubectl -n demo logs app-pod -c app --timestamps --tail=100Step 3
Compare the executed command with the intended service
Pod command and args override image defaults. Look for an accidental migration command, shell ampersand, interactive command or a missing foreground flag. Exit zero after a few milliseconds strongly suggests a lifecycle mismatch.
kubectl -n demo get pod app-pod -o jsonpath='{range .spec.containers[*]}{.name}{"\t"}{.command}{"\t"}{.args}{"\n"}{end}'Step 4
Correlate the first termination with rollout changes
Read rollout history and compare image digest, environment references, resource limits and probe definitions with the prior revision. Avoid dumping secret values while collecting configuration evidence. Reproduce with the same image and limits in a staging namespace rather than modifying a production pod’s command.
The fix
Follow the earliest causal event. Repair invalid configuration when the application exits itself; fix memory pressure when the runtime records OOMKilled; add an appropriate startup probe when liveness interrupts legitimate initialisation. Treating all three as generic restart failures wastes the strongest evidence.
Run the service in the foreground as the container’s long-lived process. Put finite setup work in an init container when it must gate startup, or in a Job when it is an independent task. A Deployment repeatedly executing a one-shot command is the wrong controller contract.
For recoverable dependency failures, use bounded connection attempts, deadlines and backoff, and report inability to serve through readiness. Keep fatal local configuration errors explicit. Infinite silent retries on a misspelt database host merely turn a crash loop into an unexplained unready pod.
If the incident began with a rollout and a known-good revision exists, restore it while preserving the failed pod’s evidence. Verify that the revision includes compatible configuration and schema assumptions; rolling an image back alone is not always a valid application rollback.
How to stop it coming back
- Test the exact image command without an interactive terminal and assert it remains running
- Retain previous-container logs centrally with image digest and configuration revision
- Distinguish startup, readiness and liveness in deployment reviews
- Alert on restart rate even when replacement containers briefly become ready
FAQ
Should I delete the pod to fix CrashLoopBackOff?
Deletion can recreate the workload but does not change its image, command or configuration. Capture evidence first. If the same deterministic fault is present, the replacement repeats it and may increase pressure on an unhealthy dependency.
Why is the exit code zero?
The command completed successfully, but its controller expected it to keep running. Check daemon mode and shell backgrounding, then decide whether the workload should be a service, init container or Job.
Is ImagePullBackOff the same thing?
No. ImagePullBackOff concerns obtaining the image before the application can start. CrashLoopBackOff follows repeated starts and terminations, so previous process logs and exit state are central to its diagnosis.
Related
Other errors engineers hit next to this one
- CROSSSLOT Keys in request don’t hash to the same slot
- MOVED and ASK replies from Redis Cluster
- Redis clients stall during KEYS on a large keyspace
- A Redis lock lease expires while the worker still runs
- command not found in a script that works interactively
- Permission denied when executing a script
- bad interpreter: No such file or directory with CRLF
- Argument list too long