Running containers in production is one thing; keeping them running reliably is another. For intermediate to advanced developers, encountering a crashing or unresponsive pod in a Kubernetes cluster is not just an annoyance—it’s a critical incident that requires swift, methodical resolution. Whether you are dealing with CrashLoopBackOff errors, image pull failures, or networking timeouts, having a systematic approach to debugging is essential.
This guide moves beyond basic syntax and dives into the practical workflow of diagnosing and fixing common Kubernetes pod issues. By the end, you will have a robust mental model for isolating failures within the container lifecycle.
Step 1: Verify Pod Status and Events
The first step in any troubleshooting session is to understand the current state of the workload. Never start by diving into logs without knowing why the pod entered its current state. The kubectl get pods command gives you the high-level status, but it is often insufficient for deep diagnostics.
Instead, always pair this with kubectl describe pod. This command provides a verbose output that includes the pod’s events, which are crucial for understanding the scheduler's actions and failure reasons.
# Check the overall status of pods in the 'default' namespace
kubectl get pods -n default
# Get detailed information about a specific pod, including events
kubectl describe pod my-app-pod-xyz123 -n default
Look closely at the Events section at the bottom of the output. Are you seeing FailedScheduling? This usually points to resource constraints (CPU/Memory requests) or node affinity issues. Are you seeing ImagePullBackOff? This indicates a registry authentication problem or a typo in the image tag. Addressing these layer-one issues is often faster than digging into application logs.
Step 2: Inspect Container Logs
Once you have confirmed the pod is in a running state (or has just crashed), the next logical step is to examine the application output. Kubernetes aggregates stdout and stderr from the container entrypoint into accessible logs.
Use kubectl logs to stream or dump the logs. If a pod is in a CrashLoopBackOff state, it means the container started, encountered an error, and exited. You must check the logs from the previous instance of the container.
# View logs for the current running instance
kubectl logs my-app-pod -n default
# View logs from the previous container instance if it crashed
kubectl logs my-app-pod --previous -n default
When analyzing these logs, look for stack traces, fatal errors, or configuration exceptions. If the logs are empty, the container might be stuck in an infinite loop or waiting for a signal, which requires further investigation via exec.
Step 3: Interactive Debugging with Exec
Sometimes, logs aren't enough. You may need to inspect the file system, check environment variables, or test network connectivity from within the container context. This is where kubectl exec becomes your most powerful tool.
# Open a bash shell in the running container
kubectl exec -it my-app-pod -- /bin/bash
# Check environment variables
printenv
# Test connectivity to an external service or database
curl -v http://internal-service:8080/health
# List files to check if config maps were mounted correctly
ls -la /etc/config
Interactive sessions are particularly useful for diagnosing volume mount issues. If your application cannot find a configuration file that exists in a ConfigMap or Secret, checking the mounted directory via exec will instantly reveal if the mount path is incorrect or if the permissions are wrong.
Step 4: Check Resource Constraints
One of the most common causes of silent pod death is hitting resource limits. If a container exceeds its CPU or memory limits, the Linux kernel’s Out-Of-Memory (OOM) killer may terminate it without providing much context in the application logs.
Check the resource usage with:
kubectl top pod my-app-pod
Compare the reported usage against the resources.limits defined in your YAML manifest. If usage is consistently near the limit, you may need to scale horizontally or adjust the limits. Additionally, check the system logs on the underlying node (via journalctl on Linux nodes) for OOM killer messages if the pod is restarting frequently.
Conclusion
Troubleshooting Kubernetes pods is less about memorizing commands and more about following a logical diagnostic path: Status → Events → Logs → Internal Inspection → Resources. By mastering these four steps, you can efficiently isolate failures, reduce mean-time-to-resolution (MTTR), and maintain the stability of your containerized applications. Remember, Kubernetes provides all the tools you need; you just need to know how to use them in sequence.