Skip to main content

Troubleshooting ImagePullBackOff and Registry Authentication Issues

Diagnose and fix ImagePullBackOff and ErrImagePull errors in Kubernetes caused by registry authentication failures, incorrect image names, and network restrictions.

52 Terms

Overview and What You Will Learn

ImagePullBackOff is one of the most common errors engineers encounter when deploying to Kubernetes — and one of the most frustrating, because the same image that pulls fine on your laptop refuses to pull on the cluster. The error can be caused by five completely different root causes that all produce the same status message. This lab walks through every cause systematically with concrete diagnostic commands and fixes.

By the end of this guide you will be able to:

  • Distinguish between ErrImagePull (first attempt) and ImagePullBackOff (retry with backoff) and why both map to the same root causes
  • Diagnose authentication failures with private registries and create the correct imagePullSecret
  • Fix image name, tag, and digest errors that cause pull failures
  • Configure registry credentials for AWS ECR, GCP Artifact Registry, and private Harbor instances
  • Resolve network-level pull failures caused by firewall rules or registry outages

Why This Matters in Production

At Razorpay, a CI/CD pipeline successfully built and pushed a new payments service image to their private ECR registry — but the Kubernetes deployment sat in ImagePullBackOff for 11 minutes before an engineer noticed. The root cause: the imagePullSecret referencing the ECR credentials had expired. The new pods could not pull the image, but the old pods were still running — so no alerts fired. The fix took 30 seconds once diagnosed, but discovery took 11 minutes of confusion.

At Hotstar, a developer accidentally deployed with image: video-encoder:latest instead of image: registry.hotstar.com/video-encoder:v3.2.1 — pulling from Docker Hub (which doesn't have the image) instead of their private registry. Same error message, completely different cause, completely different fix.

Core Principles

The five root causes of ImagePullBackOff — always check in this order:

CAUSE 1 — Wrong image name or tag

YAML
image: my-app:v2 ← no registry prefix = pulls from Docker Hub
image: registry.razorpay.in/my-app ← missing tag = tries "latest" which may not exist
image: registry.razorpay.in/my-app:v2.0.0 ← tag doesn't exist in registry

CAUSE 2 — Missing or incorrect imagePullSecret

Private registry requires credentials. No imagePullSecret = 401 Unauthorized from registry. Wrong secret = 403 Forbidden or "incorrect username or password"

CAUSE 3 — Expired credentials (ECR tokens expire every 12 hours)

AWS ECR tokens are time-limited. A secret created yesterday may be expired today. Symptom: worked before, suddenly fails on new pod scheduling.

CAUSE 4 — Network cannot reach the registry

Node firewall blocks outbound HTTPS to registry domain. Private registry behind VPN that cluster nodes cannot access. Symptom: "dial tcp: connection timed out" in pod events.

CAUSE 5 — Registry rate limiting (Docker Hub)

Docker Hub limits unauthenticated pulls to 100/6hr per IP. Shared NAT gateway = all nodes share one IP = rate limit hit quickly. Symptom: "toomanyrequests: You have reached your pull rate limit"

Detailed Step-by-Step Practical Lab

Step 1 — Identify the Exact Error

The pod status shows ImagePullBackOff:

Bash
kubectl get pods -n production
TEXT
NAME READY STATUS RESTARTS
payments-api-6d8f9b-xkp2q 0/1 ImagePullBackOff 0

ALWAYS describe the pod first — the Events section contains the actual error:

Bash
kubectl describe pod payments-api-6d8f9b-xkp2q -n production

Look for the Events section at the bottom:

TEXT
Events:
Warning Failed kubelet Failed to pull image "registry.razorpay.in/payments-api:v2.1.0":
rpc error: code = Unknown desc = failed to pull and unpack image
"registry.razorpay.in/payments-api:v2.1.0":
unexpected status code 401 Unauthorized
Warning Failed kubelet Error: ErrImagePull
Normal BackOff kubelet Back-off pulling image "registry.razorpay.in/payments-api:v2.1.0"
Warning Failed kubelet Error: ImagePullBackOff
Remember

ErrImagePull is the first failed attempt. ImagePullBackOff is Kubernetes applying exponential backoff (10s → 20s → 40s → ... → 5min cap) before retrying. Both share the same root cause — always read the Failed to pull image line above them for the actual error message.

Step 2 — Diagnose and Fix: Wrong Image Name or Tag

Error signature: "manifest unknown" or "not found"

TEXT
Failed to pull image "my-app:v2": manifest unknown: manifest tagged by "v2" is not found

Check what tags actually exist in your registry. For AWS ECR:

Bash
aws ecr describe-images \
--repository-name payments-api \
--region ap-south-1 \
--query 'imageDetails[*].imageTags' \
--output table

For Harbor (private registry):

Bash
curl -u rahul:password \
https://registry.razorpay.in/v2/payments-api/tags/list

Fix: Update the deployment with the correct image reference:

Bash
kubectl set image deployment/payments-api \
payments-api=registry.razorpay.in/payments-api:v2.1.0 \
-n production

Verify the image reference in the deployment spec:

Bash
kubectl get deployment payments-api -n production \
-o jsonpath='{.spec.template.spec.containers[0].image}'
TEXT
registry.razorpay.in/payments-api:v2.1.0
Security

Never use image: myapp:latest in production manifests. The latest tag is mutable — the registry can silently replace it with a different image. Always pin to an immutable tag (v2.1.0) or image digest (sha256:abc123...) for reproducible deployments.

Step 3 — Diagnose and Fix: Missing imagePullSecret for Private Registry

Error signature: "401 Unauthorized" or "403 Forbidden"

TEXT
Failed to pull image: unexpected status code 401 Unauthorized

Step 1 — Create the imagePullSecret from registry credentials.

Method A: Docker config file (most portable):

Bash
kubectl create secret docker-registry razorpay-registry-secret \
--docker-server=registry.razorpay.in \
--docker-username=deploy-bot \
--docker-password=sup3rs3cr3tP@ssword \
--docker-email=devops@razorpay.com \
--namespace=production

Method B: From an existing Docker config.json (if you've already logged in locally):

Bash
kubectl create secret generic razorpay-registry-secret \
--from-file=.dockerconfigjson=$HOME/.docker/config.json \
--type=kubernetes.io/dockerconfigjson \
--namespace=production

Verify the secret was created correctly:

Bash
kubectl get secret razorpay-registry-secret -n production -o yaml
YAML
apiVersion: apps/v1
kind: Deployment
metadata:
name: payments-api
namespace: production
spec:
template:
spec:
imagePullSecrets:
- name: razorpay-registry-secret # Reference secret by name
containers:
- name: payments-api
image: registry.razorpay.in/payments-api:v2.1.0
Bash
kubectl apply -f deployment-with-pull-secret.yaml

Alternative: Patch an existing deployment to add imagePullSecrets:

Bash
kubectl patch deployment payments-api -n production \
--type='json' \
-p='[{"op":"add","path":"/spec/template/spec/imagePullSecrets","value":[{"name":"razorpay-registry-secret"}]}]'
Tip

Attach the imagePullSecret to the namespace's default ServiceAccount so every pod in the namespace automatically inherits it — eliminating the need to add imagePullSecrets to every individual deployment:

kubectl patch serviceaccount default -n production -p '{"imagePullSecrets": [{"name": "razorpay-registry-secret"}]}'

Step 4 — Diagnose and Fix: Expired AWS ECR Credentials

AWS ECR authentication tokens expire every 12 hours. A secret created during cluster setup will fail the next day:

Error signature: "no basic auth credentials" or "401" from ECR

TEXT
Failed to pull image: pull access denied, repository does not exist or may require authorization

Check when the ECR secret was last updated:

Bash
kubectl get secret ecr-registry-secret -n production \
-o jsonpath='{.metadata.creationTimestamp}'
◈ DIAGRAM
2025-05-24T08:15:00Z ← created >12 hours ago = expired

Refresh the ECR token and update the secret:

Bash
aws ecr get-login-password \
--region ap-south-1 | \
kubectl create secret docker-registry ecr-registry-secret \
--docker-server=123456789.dkr.ecr.ap-south-1.amazonaws.com \
--docker-username=AWS \
--docker-password=$(aws ecr get-login-password --region ap-south-1) \
--namespace=production \
--dry-run=client -o yaml | kubectl apply -f -
YAML
# ecr-token-refresher-cronjob.yaml — automatically refresh ECR token every 6 hours
apiVersion: batch/v1
kind: CronJob
metadata:
name: ecr-token-refresher
namespace: production
spec:
schedule: "0 */6 * * *" # Every 6 hours — well within the 12-hour expiry
jobTemplate:
spec:
template:
spec:
serviceAccountName: ecr-refresher-sa # Needs IAM role to call ECR
restartPolicy: OnFailure
containers:
- name: ecr-refresher
image: amazon/aws-cli:latest
command:
- /bin/sh
- -c
- |
ECR_TOKEN=$(aws ecr get-login-password --region ap-south-1)
kubectl create secret docker-registry ecr-registry-secret \
--docker-server=123456789.dkr.ecr.ap-south-1.amazonaws.com \
--docker-username=AWS \
--docker-password=${ECR_TOKEN} \
--namespace=production \
--dry-run=client -o yaml | kubectl apply -f -
echo "ECR token refreshed at $(date)"
Remember

The permanent solution for ECR on EKS is to use IRSA (IAM Roles for Service Accounts) instead of static credentials. With IRSA, the node's IAM role automatically authorises ECR pulls with no secrets required — no tokens to expire, no CronJob refresh needed.

Step 5 — Diagnose and Fix: Network Cannot Reach Registry

Error signature: "connection timed out" or "no such host"

TEXT
Failed to pull image: dial tcp: lookup registry.razorpay.in: no such host
Failed to pull image: dial tcp 10.20.30.40:443: i/o timeout

Test DNS resolution for the registry from inside a pod on the same node:

Bash
kubectl run registry-test \
--image=busybox:1.35 \
--restart=Never \
-n production \
-- nslookup registry.razorpay.in

Test TCP connectivity to the registry port:

Bash
kubectl run registry-test-2 \
--image=nicolaka/netshoot \
--restart=Never \
-n production \
-- nc -zv registry.razorpay.in 443
◈ DIAGRAM
Connection to registry.razorpay.in 443 port [tcp/https] succeeded! ← reachable
nc: connect to registry.razorpay.in port 443 (tcp) failed: Connection timed out ← blocked

If connection times out — check node security group / firewall rules. For AWS EKS — verify the node security group allows outbound HTTPS (443) to registry IP:

Bash
aws ec2 describe-security-groups \
--group-ids sg-node-security-group-id \
--query 'SecurityGroups[0].IpPermissionsEgress'

For private registries — check if the registry is accessible from the VPC. Test directly from the node via SSH:

Bash
ssh ec2-user@mumbai-worker-node-ip
curl -v https://registry.razorpay.in/v2/
Bash
Should return: {"errors":[{"code":"UNAUTHORIZED",...}]} ← reachable (auth error is expected)
Or: curl: (6) Could not resolve host ← DNS failure
Or: curl: (28) Operation timed out ← network blocked

Clean up test pods:

Bash
kubectl delete pod registry-test registry-test-2 -n production
Step 6 — Diagnose and Fix: Docker Hub Rate Limiting

Error signature: "toomanyrequests"

TEXT
Failed to pull image: toomanyrequests:
You have reached your pull rate limit. You may increase the limit by authenticating.

Check current rate limit status from inside a pod:

Bash
kubectl run ratelimit-test \
--image=nicolaka/netshoot \
--restart=Never \
-n production \
-- sh -c "
TOKEN=\$(curl -s 'https://auth.docker.io/token?service=registry.docker.io&scope=repository:ratelimitpreview/test:pull' | jq -r .token)
curl -s --head -H \"Authorization: Bearer \$TOKEN\" https://registry-1.docker.io/v2/ratelimitpreview/test/manifests/latest 2>&1 | grep -i ratelimit
"
◈ DIAGRAM
ratelimit-limit: 100;w=21600
ratelimit-remaining: 0;w=21600 ← exhausted

Fix 1: Authenticate Docker Hub pulls to get higher limits (200/6hr per account). Create a Docker Hub pull secret:

Bash
kubectl create secret docker-registry dockerhub-secret \
--docker-server=https://index.docker.io/v1/ \
--docker-username=razorpay-devops \
--docker-password=dckr_pat_xxxxxxxxxxxx \
--namespace=production

Fix 2 (Permanent): Mirror public images to your private registry. Never pull from Docker Hub directly in production — mirror images first.

Mirror a public image to your private ECR registry — pull locally, retag, push to private registry:

Bash
docker pull postgres:15.4
docker tag postgres:15.4 123456789.dkr.ecr.ap-south-1.amazonaws.com/postgres:15.4
docker push 123456789.dkr.ecr.ap-south-1.amazonaws.com/postgres:15.4

Update deployments to use the mirrored image:

Bash
kubectl set image statefulset/postgres \
postgres=123456789.dkr.ecr.ap-south-1.amazonaws.com/postgres:15.4 \
-n production
Step 7 — Verify the Fix and Confirm Successful Pull

After applying any fix — force a new pod to attempt the pull:

Bash
kubectl rollout restart deployment/payments-api -n production

Watch the new pod status:

Bash
kubectl get pods -n production -w
◈ DIAGRAM
payments-api-7f8g9h-mn3lp 0/1 ContainerCreating 0 5s
payments-api-7f8g9h-mn3lp 1/1 Running 0 18s ← image pulled successfully

Confirm image was pulled by checking pod events:

Bash
kubectl describe pod payments-api-7f8g9h-mn3lp -n production | grep -A5 Events
◈ DIAGRAM
Events:
Normal Pulling kubelet Pulling image "registry.razorpay.in/payments-api:v2.1.0"
Normal Pulled kubelet Successfully pulled image in 4.821s ← success
Normal Created kubelet Created container payments-api
Normal Started kubelet Started container payments-api

Verify which image digest was actually pulled:

Bash
kubectl get pod payments-api-7f8g9h-mn3lp -n production \
-o jsonpath='{.status.containerStatuses[0].imageID}'
Bash
docker-pullable://registry.razorpay.in/payments-api@sha256:abc123def456...

Production Best Practices & Common Pitfalls

  • Mirror all public images (Docker Hub, quay.io, gcr.io) to your private registry as part of your base image policy. Public registries have rate limits, availability incidents, and can remove images — your production cluster should never depend on them directly.
  • Use image digests (image: registry.razorpay.in/payments-api@sha256:abc123...) instead of mutable tags in production GitOps manifests. Tags can be overwritten; digests are immutable.
  • Rotate registry credentials on a schedule and automate the Kubernetes secret update via CI/CD or a CronJob — manual rotation always gets forgotten until a deployment fails at 2am.
  • For multi-namespace clusters, attach the imagePullSecret to each namespace's default ServiceAccount rather than adding it to every deployment manifest individually. One change, universal coverage.
  • Always test image pull independently of deployment configuration by running kubectl run test --image=<your-image> --restart=Never -n <ns> — this isolates the pull failure from any deployment spec issues.
Common Mistake

Deleting and recreating the pod to "force a retry" of an ImagePullBackOff. The backoff timer resets on pod recreation, but if the root cause is not fixed (wrong image name, missing secret, expired token), the new pod will fail identically. Fix the root cause first, confirmed by kubectl describe pod, before attempting any restart.

Quick Reference & Troubleshooting Commands

Command Purpose
kubectl describe pod <name> -n <ns> Primary diagnostic — read the Events section for the exact error
kubectl get events -n <ns> --field-selector reason=Failed List all pull failure events in the namespace
kubectl get secret <name> -n <ns> -o yaml Inspect imagePullSecret contents
kubectl create secret docker-registry <name> --docker-server=... --docker-username=... --docker-password=... Create registry pull secret
kubectl patch serviceaccount default -n <ns> -p '{"imagePullSecrets": [{"name": "<secret>"}]}' Attach pull secret to all pods in namespace
kubectl set image deployment/<name> <container>=<new-image> -n <ns> Fix image reference directly
aws ecr get-login-password --region <region> Generate fresh ECR auth token
kubectl run test --image=<image> --restart=Never -n <ns> Test image pull in isolation
kubectl rollout restart deployment/<name> -n <ns> Force new pods after fixing the root cause
kubectl get pod <name> -n <ns> -o jsonpath='{.status.containerStatuses[0].imageID}' Confirm which image digest was pulled

Resources

Velero vs etcd Snapshot: Not a Real Choice

Velero vs etcd Snapshot: Not a Real Choice

Velero and etcd snapshots protect different layers of a cluster, not the same thing. Here's what each covers and why production DR needs both.

5 min read•Aug 2026
Istio Ambient vs Linkerd in 2026

Istio Ambient vs Linkerd in 2026

Istio Ambient killed the sidecar-tax argument. The real 2026 decision is waypoint topology and Buoyant's licensing shift, not features vs simplicity.

5 min read•Aug 2026
Nginx Ingress vs Traefik vs Gateway API in 2026

Nginx Ingress vs Traefik vs Gateway API in 2026

Ingress-nginx retired in March 2026. Here's how Traefik and the Gateway API actually compare as replacements — and why "just swap it" is the wrong frame.

5 min read•Aug 2026
OPA Gatekeeper vs Kyverno: Policy Engine in 2026

OPA Gatekeeper vs Kyverno: Policy Engine in 2026

OPA Gatekeeper vs Kyverno compared for 2026 - Rego vs YAML, mutation maturity, operational overhead, and which policy engine fits your cluster.

5 min read•Aug 2026
Helm vs Kustomize in 2026: Templating vs Patching

Helm vs Kustomize in 2026: Templating vs Patching

Helm vs Kustomize compared for 2026 - templating vs patching, Helm 4's new features, and why most production teams end up running both.

5 min read•Aug 2026
Prometheus vs Datadog vs New Relic: Real Costs

Prometheus vs Datadog vs New Relic: Real Costs

Prometheus, Datadog, and New Relic compared for 2026 - real pricing at scale, hidden cost drivers, and which fits a Kubernetes-heavy stack.

5 min read•Aug 2026
Cluster Autoscaler vs Karpenter for EKS in 2026

Cluster Autoscaler vs Karpenter for EKS in 2026

Cluster Autoscaler vs Karpenter compared for EKS in 2026 - provisioning speed, bin-packing, cloud support, and when each is the right default.

5 min read•Aug 2026
GKE vs EKS vs AKS in 2026: Which Fits Your Team?

GKE vs EKS vs AKS in 2026: Which Fits Your Team?

GKE, EKS, and AKS compared for 2026 - control plane pricing, Autopilot vs Karpenter vs Node Auto Provisioning, and which platform actually fits your team.

5 min read•Aug 2026
K3s vs K8s vs MicroK8s in 2026

K3s vs K8s vs MicroK8s in 2026

K3s, full Kubernetes, and MicroK8s compared for 2026 - resource footprint, production readiness, and which fits edge, homelab, or cloud workloads.

5 min read•Aug 2026
Canary Deployments with Argo Rollouts & Flagger

Canary Deployments with Argo Rollouts & Flagger

Ship to 5% of users first and auto-rollback in minutes — a hands-on guide to canary deployments with Argo Rollouts and Flagger on Kubernetes.

5 min read•Jun 2026
OpenTelemetry Explained: Metrics, Logs, Traces

OpenTelemetry Explained: Metrics, Logs, Traces

OpenTelemetry unifies metrics, logs, and traces under one open standard — how it works, what it replaces, and how to instrument a service in 20 minutes.

5 min read•Jun 2026
Kubernetes Cost Optimization Without Breaking SLOs

Kubernetes Cost Optimization Without Breaking SLOs

Average Kubernetes CPU utilization across production clusters is 8%. Here is the complete 2026 playbook for cutting cloud spend without touching your SLOs.

5 min read•Jun 2026
ArgoCD vs FluxCD: GitOps for Kubernetes in 2026

ArgoCD vs FluxCD: GitOps for Kubernetes in 2026

ArgoCD and FluxCD are the two dominant GitOps engines for Kubernetes in 2026 — this breakdown tells you exactly which one to pick and why.

10 min read•Jun 2026

Explore More in Kubernetes Observability and Scaling

All 3 Topics

Frequently Asked Questions

Is Troubleshooting ImagePullBackOff and Registry Authentication Issues free to learn on DevOps Network?

Yes - this topic, like everything on DevOps Network, is 100% free with no paywall or sign-up gate.

What does the Troubleshooting ImagePullBackOff and Registry Authentication Issues topic cover?

Diagnose and fix ImagePullBackOff and ErrImagePull errors in Kubernetes caused by registry authentication failures, incorrect image names, and network restrictions.