Troubleshooting — The Manager's Problem Desk
Everyone eventually sees a scary Kubernetes error. This is the Problem Desk: what the message means, why it happened, and how to get out. Each problem follows the same Problem → Cause → Fix rhythm you know.
1 · CrashLoopBackOff
Pod status shows CrashLoopBackOff and keeps restarting.
Cause: The container starts, crashes, and Kubernetes restarts it. This repeats with exponential backoff. Usually a bug in the app, a missing file, or a wrong command.
Fix:
kubectl logs <pod-name>
kubectl logs --previous <pod-name>
kubectl describe pod <pod-name>
CrashLoopBackOff = the app crashes on startup. The logs always tell you why. Fix the app, rebuild, reapply.
2 · ImagePullBackOff
Pod status shows ImagePullBackOff or ErrImagePull.
Cause: Kubernetes can't pull the container image. Wrong name, wrong tag, or the registry requires authentication.
Fix:
kubectl describe pod <pod-name> | grep Image
docker pull <image-name>
kubectl create secret docker-registry regcred \
--docker-server=registry.example.com \
--docker-username=user \
--docker-password=pass
Image names are case-sensitive. nginx ≠ Nginx. Always check the exact name and tag.
3 · Pod stuck in Pending
Pod status shows Pending and never starts.
Cause: Kubernetes can't find a node with enough resources, or the node selector/affinity rules don't match any node.
Fix:
kubectl describe pod <pod-name>
Look for Events — it will say "Insufficient cpu" or "no nodes match selector."
kubectl top nodes
# Edit the deployment and lower cpu/memory requests
Pending means Kubernetes wants to schedule the Pod but can't. The describe output tells you exactly why.
4 · Service can't reach the Pod
The Service has no endpoints, or curling it returns nothing.
Cause: The Service's selector doesn't match any Pod labels, or the Pods aren't Ready.
Fix:
kubectl get endpoints <service-name>
If endpoints are empty, the selector is wrong.
kubectl get pods --show-labels
kubectl get pods -o wide
A Service only sends traffic to Ready Pods. If readiness probes fail, the Pod is excluded from the Service.
5 · OOMKilled
Pod restarts with reason OOMKilled.
Cause: The container exceeded its memory limit and the OS killed it.
Fix:
kubectl top pod <pod-name>
# Edit the deployment:
kubectl edit deployment <deployment-name>
# Change resources.limits.memory to a higher value
OOMKilled means the app uses more memory than you allowed. Increase the limit or fix the memory leak.
6 · "connection refused" when accessing a Service
curl to a Service returns "connection refused."
Cause: The Pod is listening on a different port than the Service expects, or the app isn't started yet.
Fix:
kubectl exec -it <pod-name> -- sh -c "netstat -tlnp"
kubectl get svc <service-name> -o yaml
Service port is what clients connect to. targetPort is what the Pod listens on. They must match.