Linux / shell
ImagePullBackOff and ErrImagePull
Written and reviewed by Sahil Srivastav
ErrImagePull
ImagePullBackOffWhat this error actually means
Kubelet asked the container runtime to obtain an image and that operation failed. ErrImagePull records the failed attempt; ImagePullBackOff means Kubernetes is waiting before retrying. Neither status tells you whether the registry rejected authentication, the reference was missing, or the node could not reach the registry at all. The detailed event does.
The network origin matters: the node’s runtime pulls images, not the application process inside the pod. A curl command that succeeds on a developer laptop or an existing pod does not prove that the node has the same DNS, proxy, firewall, certificate trust or credentials.
A cached image can disguise an invalid deployment configuration. Existing nodes start successfully using a local copy while a newly autoscaled node fails to pull it. Rebuilding the same mutable tag adds another ambiguity because two machines may then run different content under one apparent version.
Causes, most common first
- 1The image reference does not resolve to published content. A wrong repository path, missing release tag or deployment racing an unfinished image push is the first check. Some registries obscure nonexistent repositories behind authorization errors, so verify both the exact reference and the identity allowed to read it.
- 2Pull credentials are absent, expired or attached in the wrong namespace. An imagePullSecrets reference resolves in the pod’s namespace. A correctly named Secret in another namespace is still unavailable. The service account used by this workload may also differ from the one whose registry access was configured.
- 3The node cannot complete the registry connection. Node DNS, egress policy, proxies or an untrusted registry certificate can stop the pull before authentication matters. Changing the application’s CA bundle cannot repair the container runtime’s trust store.
- 4The registry rejects demand or lacks the required platform. Rate limiting can surface during a large rollout, while a single-platform image fails on a different CPU architecture. The latter usually produces a manifest-selection error before execution; it is distinct from a bad binary that starts and returns exec format error.
When you see it
- The container has no application logs because its process never started
- A rollout fails only on new nodes or in one availability zone
- Events include unauthorized, manifest unknown, certificate, DNS or timeout details
- Older replicas continue serving while replacements remain waiting
How to diagnose it
Step 1
Copy the full pull event, not just the status column
Substitute the real pod and namespace for app-pod and demo. The text following Failed to pull image identifies the registry response or transport failure. Record the node name to compare failing and healthy placements.
kubectl -n demo describe pod app-pod
kubectl -n demo get pod app-pod -o wideStep 2
Inspect image references and credential references safely
Check the live pod because a service account or admission controller may supply pull-secret references. Inspect the Secret’s existence and type without printing its credential data into terminal logs.
kubectl -n demo get pod app-pod -o jsonpath='{.spec.containers[*].image}{"\n"}{.spec.serviceAccountName}{"\n"}{.spec.imagePullSecrets[*].name}{"\n"}'
kubectl -n demo get secret registry-credentials -o jsonpath='{.type}{"\n"}'Step 3
Check the published platform manifest
With authorised registry access, inspect the same immutable image reference the deployment uses. Confirm that the manifest contains the node architecture. Laptop success can be misleading when the laptop runs arm64 and the cluster uses amd64, or the reverse.
docker buildx imagetools inspect registry.example.com/team/app:release-42
kubectl get nodes -L kubernetes.io/archStep 4
Test from the failing infrastructure boundary
Use approved node diagnostics or runtime logs to verify registry DNS, TLS and egress from that node. Do not copy production credentials into an ad hoc command line. If only cold nodes fail, compare registry access with nodes using cached images before changing the deployment.
The fix
Publish and verify the release image before updating the workload. Prefer immutable digests for deployments so the reference identifies the content actually tested. If a tag was mistyped, correct the workload manifest rather than repeatedly deleting waiting pods.
Provision registry credentials through the existing secret-management path in the pod’s namespace and attach them to the intended pod or service account. Confirm repository read permission and token expiry. Registry credentials are separate from the application’s runtime credentials.
Repair the node’s DNS, egress route or runtime certificate trust when the event shows a transport failure. Disabling TLS verification creates a different problem and does not explain why this node’s trusted registry configuration drifted.
For missing platform support, publish a correctly built image for the node architecture or a multi-platform index containing validated variants. For throttling, reduce rollout concurrency and use a registry cache or mirror with an explicit freshness policy. More simultaneous retries can extend the outage.
How to stop it coming back
- Validate that an image digest is pullable before advancing the deployment stage
- Exercise cold-node pulls so cache hits cannot hide missing permissions
- Monitor registry token expiry and image-pull failures by node and repository
- Keep platform requirements and published manifest variants aligned in CI
FAQ
Will kubectl logs explain the failure?
Usually not: the process has not started. Read pod events and node runtime diagnostics. Application logging cannot describe a registry operation performed before the container exists.
Can I set imagePullPolicy to Never?
Only for an environment that deliberately guarantees the image on every eligible node. As an incident workaround it replaces a registry dependency with an undocumented cache dependency and fails as soon as scheduling selects a cold node.
Why does docker pull succeed on my laptop?
It uses your laptop’s credentials, network, certificate store and selected platform. The Kubernetes node can differ in every one of those dimensions, so local success narrows the problem without proving node access.
Related
Other errors engineers hit next to this one
- ThreadLocal value leaking across requests on a pooled thread
- Consumer stuck in Object.wait() with work already in the queue
- java.lang.IllegalMonitorStateException: current thread is not owner
- Thread pool starvation — every worker waiting on a task in its own pool
- Partially constructed object published by double-checked locking
- Lost update from get-then-put on a ConcurrentHashMap
- CompletableFuture failed with nothing logged
- awaitTermination never returns and the JVM will not exit