Linux / shell

CrashLoopBackOff: diagnose the exit behind the backoff

Written and reviewed by Sahil Srivastav

KubernetesContainer lifecycleRestart loops
CrashLoopBackOff

What this error actually means

CrashLoopBackOff describes a waiting interval between repeated container restarts. It is a kubelet behaviour, not the exception that broke the application. The important evidence sits one level earlier: the terminated process’s exit status, its previous logs, or an event showing that a probe asked kubelet to kill it.

The word crash is misleading when the command exits successfully. A Deployment normally expects a long-running process and uses restartPolicy Always. A migration command that completes with exit code zero, or a server that daemonises and lets its original process exit, can be restarted repeatedly even though it never threw an exception.

Backoff protects the node and dependencies from an uncontrolled restart storm. Its timing can vary with Kubernetes configuration and version; do not infer a root cause from the precise delay. Deleting the pod may reset the visible cycle while throwing away the record you needed to diagnose it.

Causes, most common first

  1. 1Application startup fails deterministically. A malformed configuration, unsupported runtime option or unavailable required dependency makes every attempt exit at the same point. Compare the first failing revision with the last working image and configuration; changing restart timing cannot make invalid input valid.
  2. 2Memory enforcement or a liveness probe terminates the process. These failures are outside the application’s normal exception path. OOMKilled implicates memory; repeated Unhealthy and Killing events implicate the probe. A shutdown stack trace can be a consequence of the kill rather than the initiating problem.
  3. 3The container command has the wrong lifecycle. A CLI finishes, a shell backgrounds the server, or the executable starts in daemon mode. Kubernetes tracks the container’s main process, not the existence of an unrelated child after that process exits.
  4. 4A dependency outage is coupled to process survival. An application exits on each failed database connection, turning a dependency incident into repeated cold starts. Bounded reconnect logic may be appropriate after startup; pretending the dependency is healthy is not. Separate serving readiness from recoverable connection loss.

When you see it

  • The restart count rises while the container alternates between Running and Waiting
  • Current logs contain only startup lines, but --previous shows the fatal event
  • A rollout stalls because replacement pods never remain ready long enough
  • The application starts normally and is killed at a repeatable probe deadline

How to diagnose it

Step 1

Inspect the affected container’s last termination

Use the actual namespace, pod and container in place of demo, app-pod and app. In a multi-container pod, the main application can be healthy while a sidecar loops. Init-container status is separate and should be inspected when the application never starts.

kubectl -n demo get pod app-pod -o jsonpath='{range .status.containerStatuses[*]}{.name}{"\t"}{.restartCount}{"\t"}{.lastState.terminated}{"\n"}{end}'
kubectl -n demo describe pod app-pod

Step 2

Read logs from the instance that exited

The new process may not have reached the failing code yet. Previous logs normally cover only the previous terminated instance, so export them promptly and use central logging for older attempts. Empty logs also matter: execution or configuration may fail before application logging initialises.

kubectl -n demo logs app-pod -c app --previous --timestamps --tail=200
kubectl -n demo logs app-pod -c app --timestamps --tail=100

Step 3

Compare the executed command with the intended service

Pod command and args override image defaults. Look for an accidental migration command, shell ampersand, interactive command or a missing foreground flag. Exit zero after a few milliseconds strongly suggests a lifecycle mismatch.

kubectl -n demo get pod app-pod -o jsonpath='{range .spec.containers[*]}{.name}{"\t"}{.command}{"\t"}{.args}{"\n"}{end}'

Step 4

Correlate the first termination with rollout changes

Read rollout history and compare image digest, environment references, resource limits and probe definitions with the prior revision. Avoid dumping secret values while collecting configuration evidence. Reproduce with the same image and limits in a staging namespace rather than modifying a production pod’s command.

The fix

Follow the earliest causal event. Repair invalid configuration when the application exits itself; fix memory pressure when the runtime records OOMKilled; add an appropriate startup probe when liveness interrupts legitimate initialisation. Treating all three as generic restart failures wastes the strongest evidence.

Run the service in the foreground as the container’s long-lived process. Put finite setup work in an init container when it must gate startup, or in a Job when it is an independent task. A Deployment repeatedly executing a one-shot command is the wrong controller contract.

For recoverable dependency failures, use bounded connection attempts, deadlines and backoff, and report inability to serve through readiness. Keep fatal local configuration errors explicit. Infinite silent retries on a misspelt database host merely turn a crash loop into an unexplained unready pod.

If the incident began with a rollout and a known-good revision exists, restore it while preserving the failed pod’s evidence. Verify that the revision includes compatible configuration and schema assumptions; rolling an image back alone is not always a valid application rollback.

How to stop it coming back

  • Test the exact image command without an interactive terminal and assert it remains running
  • Retain previous-container logs centrally with image digest and configuration revision
  • Distinguish startup, readiness and liveness in deployment reviews
  • Alert on restart rate even when replacement containers briefly become ready

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Should I delete the pod to fix CrashLoopBackOff?

Deletion can recreate the workload but does not change its image, command or configuration. Capture evidence first. If the same deterministic fault is present, the replacement repeats it and may increase pressure on an unhealthy dependency.

Why is the exit code zero?

The command completed successfully, but its controller expected it to keep running. Check daemon mode and shell backgrounding, then decide whether the workload should be a service, init container or Job.

Is ImagePullBackOff the same thing?

No. ImagePullBackOff concerns obtaining the image before the application can start. CrashLoopBackOff follows repeated starts and terminations, so previous process logs and exit state are central to its diagnosis.

Related

Other errors engineers hit next to this one

Full error and symptom index →