Skip to main content

Troubleshooting — The Manager's Problem Desk

Everyone eventually sees a scary Kubernetes error. This is the Problem Desk: what the message means, why it happened, and how to get out. Each problem follows the same Problem → Cause → Fix rhythm you know.

1 · CrashLoopBackOff​

Problem

Pod status shows CrashLoopBackOff and keeps restarting.

Cause: The container starts, crashes, and Kubernetes restarts it. This repeats with exponential backoff. Usually a bug in the app, a missing file, or a wrong command.

Fix:

Check the logs
kubectl logs <pod-name>
Check previous container logs (if it crashed before logging)
kubectl logs --previous <pod-name>
Describe the pod for events
kubectl describe pod <pod-name>
Remember

CrashLoopBackOff = the app crashes on startup. The logs always tell you why. Fix the app, rebuild, reapply.

2 · ImagePullBackOff​

Problem

Pod status shows ImagePullBackOff or ErrImagePull.

Cause: Kubernetes can't pull the container image. Wrong name, wrong tag, or the registry requires authentication.

Fix:

Check the image name in the Deployment
kubectl describe pod <pod-name> | grep Image
Test pulling manually
docker pull <image-name>
If it's a private registry, create a secret
kubectl create secret docker-registry regcred \
--docker-server=registry.example.com \
--docker-username=user \
--docker-password=pass
Remember

Image names are case-sensitive. nginx ≠ Nginx. Always check the exact name and tag.

3 · Pod stuck in Pending​

Problem

Pod status shows Pending and never starts.

Cause: Kubernetes can't find a node with enough resources, or the node selector/affinity rules don't match any node.

Fix:

Check why it's pending
kubectl describe pod <pod-name>

Look for Events — it will say "Insufficient cpu" or "no nodes match selector."

Check node resources
kubectl top nodes
Reduce resource requests if needed
# Edit the deployment and lower cpu/memory requests
Remember

Pending means Kubernetes wants to schedule the Pod but can't. The describe output tells you exactly why.

4 · Service can't reach the Pod​

Problem

The Service has no endpoints, or curling it returns nothing.

Cause: The Service's selector doesn't match any Pod labels, or the Pods aren't Ready.

Fix:

Check Service endpoints
kubectl get endpoints <service-name>

If endpoints are empty, the selector is wrong.

Check Pod labels match the Service selector
kubectl get pods --show-labels
Check Pod readiness
kubectl get pods -o wide
Remember

A Service only sends traffic to Ready Pods. If readiness probes fail, the Pod is excluded from the Service.

5 · OOMKilled​

Problem

Pod restarts with reason OOMKilled.

Cause: The container exceeded its memory limit and the OS killed it.

Fix:

Check memory usage
kubectl top pod <pod-name>
Increase the memory limit
# Edit the deployment:
kubectl edit deployment <deployment-name>
# Change resources.limits.memory to a higher value
Remember

OOMKilled means the app uses more memory than you allowed. Increase the limit or fix the memory leak.

6 · "connection refused" when accessing a Service​

Problem

curl to a Service returns "connection refused."

Cause: The Pod is listening on a different port than the Service expects, or the app isn't started yet.

Fix:

Check which port the Pod listens on
kubectl exec -it <pod-name> -- sh -c "netstat -tlnp"
Make sure Service targetPort matches
kubectl get svc <service-name> -o yaml
Remember

Service port is what clients connect to. targetPort is what the Pod listens on. They must match.