INSIGHTS
DevOps & Automation

CNCF CKA: Troubleshooting Kubernetes Nodes and Workloads

In this article
  1. Begin with scope and recent change history
  2. Debug Pending Pods before looking at container logs
  3. Investigate CrashLoopBackOff with current and previous logs
  4. Troubleshoot readiness, liveness, and startup probes independently
  5. Use logs, exec, and debug containers according to the situation
  6. Investigate Node NotReady as a node and control-plane communication problem
  7. Troubleshoot Services and DNS from endpoint to client
  8. Trace storage problems through PVC, PV, driver, and node
  9. Finish incidents with a verified fix and evidence-based prevention

Kubernetes troubleshooting is fastest when administrators identify which layer owns the symptom before changing anything. A failing application might be caused by its container, Pod configuration, scheduling, storage, networking, DNS, a node, or the control plane. Kubernetes exposes status, events, logs, and object relationships that let operators narrow the problem methodically.

The current CKA gives troubleshooting 30% of the exam, the largest single domain. That reflects real administration: creating a resource is only the beginning. Operators must interpret Pending Pods, restarts, failed mounts, node conditions, Service problems, and application symptoms without destroying useful evidence.

Begin with scope and recent change history

First determine whether the problem affects one container, one Pod, one Deployment, one node, one namespace, or the entire cluster. Scope immediately narrows likely causes. One failing replica suggests a local Pod or node issue; all replicas failing after a release suggest a shared application or configuration problem.

Ask what changed. New images, ConfigMaps, Secrets, RBAC, NetworkPolicy, node upgrades, CNI changes, storage maintenance, or certificate rotations are often more informative than the first error message.

A disciplined troubleshooting process preserves evidence. Restarting everything may temporarily hide the symptom while removing the logs, events, or failed state needed to explain it.

Start with concise status views. Pod phase, readiness, restart counts, node assignment, and age often reveal the first branch. Then use describe to inspect conditions and events. Kubernetes events commonly explain scheduling failures, image-pull errors, probe failures, volume problems, and node conditions.

Events are time-limited clues, not a permanent audit history. Capture important information during an incident rather than assuming it will remain available indefinitely.

Compare the child and owner objects. A failing Pod created by a Deployment may be replaced automatically; inspect the Deployment and ReplicaSet to see whether the failure is systemic to the template.

Debug Pending Pods before looking at container logs

A Pending Pod may not have started any container, so there may be no useful application log yet. Inspect scheduling events for insufficient CPU or memory, affinity mismatch, untolerated taints, unbound PVCs, port conflicts, or other placement constraints.

Check resource requests against node allocatable capacity. Verify node labels, taints, topology, and storage binding. If the scheduler says that no node matches a required condition, restarting the Pod does not change the underlying constraint.

Scheduling is a decision problem. Fix the requirement, capacity, or label state that makes every node ineligible.

Investigate CrashLoopBackOff with current and previous logs

CrashLoopBackOff means a container repeatedly starts and exits, with Kubernetes applying backoff between restarts. Inspect container logs and, when a container has restarted, previous logs so the output from the last failed instance is not lost.

Then inspect termination reason and exit code. Common causes include invalid command arguments, missing configuration, failed dependency connections, permissions, out-of-memory termination, or application startup bugs.

Do not automatically increase restart limits or disable probes. The backoff is a symptom of repeated failure, not the root cause.

ImagePullBackOff or ErrImagePull occurs before the application container runs. Check the image name and tag, registry availability, network path, authentication, imagePullSecrets, and whether the requested platform architecture is available.

A successful image pull followed by a crash is a different problem. In that case, inspect container startup, command, environment, mounted configuration, and application logs.

Supply-chain troubleshooting should also ask whether the image digest is the one intended for the release. Mutable tags can make two clusters pull different image contents under the same human-readable tag.

Troubleshoot readiness, liveness, and startup probes independently

A Pod can be Running but not Ready. That often means the container process exists but the readiness probe is failing, so Services should not send normal traffic to it. Test the probe endpoint from the Pod context and confirm port, path, protocol, and timing.

Liveness failures cause restarts. If a liveness probe depends on a remote database or another service, a downstream outage can restart healthy application processes repeatedly and amplify the incident. Liveness should normally measure whether the local process is irrecoverably unhealthy.

Startup probes are useful for slow initialization so liveness does not kill the application before it has finished starting.

Use logs, exec, and debug containers according to the situation

Container logs are usually the first application-level evidence. exec can inspect a running container when the image includes useful tools and the security policy permits it. Ephemeral debug containers can provide troubleshooting tools without rebuilding the application image.

Kubernetes also supports node debugging with kubectl debug, which can create a troubleshooting Pod on a node and expose the node filesystem under a mounted path. This is valuable when SSH is unavailable or intentionally disabled.

The general method in Linux troubleshooting still applies: observe before changing, compare healthy and unhealthy systems, and narrow the failing subsystem.

Investigate Node NotReady as a node and control-plane communication problem

When a node becomes NotReady, inspect node conditions and events. Common areas include kubelet health, container runtime status, disk or memory pressure, network connectivity to the API server, certificates, and node filesystem capacity.

If only one node is affected, compare it with healthy nodes: kubelet version, runtime, system services, CNI state, disk usage, routes, and recent maintenance. If many nodes become NotReady simultaneously, look for shared control-plane or network dependencies.

Node problems can cascade into workload symptoms as controllers replace or reschedule Pods, so timeline analysis matters.

Troubleshoot Services and DNS from endpoint to client

For Service failures, start with backend Pod readiness and EndpointSlices. Confirm the selector and ports. Test direct Pod connectivity when appropriate, then the Service IP, then the Service DNS name. This separates endpoint, Service-proxy, and DNS problems.

If DNS fails, inspect CoreDNS or the cluster’s DNS implementation. A Service can be healthy while name resolution is broken. If Service access works internally but not externally, move outward to Ingress, Gateway, load balancer, firewall, and public DNS layers.

Infrastructure logs can complement application evidence. The principles in network monitoring logs are useful when packet paths or external network devices participate in the failure.

Trace storage problems through PVC, PV, driver, and node

A Pod waiting for storage often exposes the problem through events. Check whether the PVC is Bound, whether the StorageClass exists, and whether provisioning succeeded. Then inspect attach or mount failures on the Pod and relevant CSI controller or node-plugin logs.

Topology can matter. A volume in one zone may not be attachable to a Pod scheduled in another. Node attachment limits, credentials, backend outages, and filesystem errors can also produce similar surface symptoms.

Do not delete claims or volumes casually during troubleshooting. Destructive storage actions can turn a recoverable scheduling problem into permanent data loss.

CPU saturation can produce latency without obvious container crashes. Memory pressure can trigger OOM kills. Disk pressure can prevent new Pods or cause image cleanup behavior. Inode exhaustion can break applications even when disk capacity looks acceptable.

Compare requests, limits, actual usage, and node capacity. A container killed for exceeding its memory limit has a different remediation from a node running out of memory because many Pods were packed too tightly.

Resource troubleshooting needs trend data. A one-time snapshot may look normal after an overload has already passed, so retain metrics long enough to correlate symptoms with resource spikes.

Finish incidents with a verified fix and evidence-based prevention

After identifying the root cause, apply the smallest safe correction and verify the full user path. A Pod becoming Running is not enough if the Service still has no ready endpoint or the application still returns errors.

Document what detected the problem, what evidence identified the cause, and what control would prevent recurrence. Better probes, safer configuration rollout, capacity alerts, RBAC, backup testing, or clearer ownership may matter more than the immediate command used during the incident.

The application-development view of Kubernetes is useful during these reviews because many incidents cross the boundary between platform and application responsibility.

Use timestamps to build an incident timeline. Compare the first application error with Pod restarts, Deployment changes, node conditions, autoscaling events, certificate rotations, and infrastructure maintenance. A timeline often separates cause from consequence: a node failure may trigger Pod restarts, which then trigger readiness errors, rather than the readiness probe being the original problem.

Namespaces provide a useful scoping tool. If the same platform component works in one namespace but not another, compare NetworkPolicy, ResourceQuota, LimitRange, service accounts, Secrets, admission policy, and namespace labels. If every namespace is affected, shared cluster infrastructure becomes more likely.

ResourceQuota and LimitRange can cause creation failures or unexpected resource defaults. A manifest that works in a lab namespace may be rejected in production because the namespace has stricter quota. Read the admission error before changing the workload itself.

OOMKilled containers require two levels of investigation. If the process exceeded its container memory limit, adjust the application’s memory behavior or the limit based on evidence. If the node is under memory pressure, look at aggregate workload placement and eviction behavior. Raising one container limit can worsen node pressure when the cluster lacks headroom.

CPU limits can produce throttling without restarting the container. High latency under load may be caused by CPU cgroup throttling even when average CPU graphs look moderate. Correlate latency with throttling metrics, requests, limits, and node utilization before assuming the application code slowed down.

Network debugging should use controlled tests. Compare Pod-to-Pod, Pod-to-Service, Service-by-DNS, and external paths. Use curl, dig, nc, tcpdump, or equivalent tools from an appropriate debug container when policy permits. Testing each hop turns a vague “network issue” into a specific failed transition.

NetworkPolicy can create asymmetric symptoms. One direction may be allowed while the return path or a required DNS flow is blocked, depending on policy and implementation. Examine ingress and egress policy together with namespace and Pod selectors, and verify that the CNI actually enforces the policy features in use.

Certificate failures often surface as generic TLS or connection errors. Check system time, certificate expiry, trust roots, server names, and rotation history. In Kubernetes, certificates may secure the API, admission webhooks, kubelet communication, ingress, service mesh traffic, or application endpoints, so identify which TLS relationship is failing before replacing certificates broadly.

Admission webhooks can block object creation when the webhook service is unhealthy or its certificate is invalid. If many unrelated workloads suddenly cannot be created or updated, inspect API admission errors and webhook health. The application images may be completely unrelated to the incident.

Use rollback as a diagnostic tool only when it is safe. If a new Deployment revision triggered the problem and the previous version is known to be compatible with current data and configuration, rollback can restore service quickly. But rolling back code after an irreversible schema migration may create a second failure, so release design should define rollback compatibility in advance.

Control-plane symptoms require a different branch from application symptoms. If many kubectl requests are slow or failing, inspect API server availability, etcd health, admission webhooks, and control-plane network paths. Existing application traffic may remain healthy while management operations degrade, so application monitoring alone can miss the incident.

Compare desired and actual state. A Deployment may request five replicas while only three are available, a node may be Ready but cordoned, or a PVC may be Bound while the mount still fails. Kubernetes exposes both declaration and status; troubleshooting is often the process of explaining why the two differ.

Use temporary debug tooling carefully in production. Privileged debug Pods, host filesystem access, or packet capture can expose credentials and sensitive traffic. Apply the same RBAC, audit, and cleanup discipline to troubleshooting tools that you would apply to normal administrative access.

Runbooks should identify safe first commands and destructive boundaries. Reading events and logs is low risk; deleting a Pod, force-deleting a namespace, removing a finalizer, or editing etcd state can have much larger consequences. Clear escalation points reduce the temptation to use destructive fixes under pressure.

Post-incident work should improve detection as well as prevention. If the root cause was obvious only after manually checking an obscure metric, add an alert or dashboard that surfaces that condition next time. Troubleshooting knowledge becomes operational maturity when the system gets easier to diagnose after every incident.

Healthy troubleshooting culture also separates mitigation from root cause. Scaling replicas, restarting a node, or rolling back a release may restore service quickly, but the incident should remain open until the team understands why the action worked. Otherwise the same failure is likely to return under similar conditions.

Keep a small library of known-good diagnostic Pods or ephemeral-container images with tools such as curl, dig, ip, ss, and tcpdump, subject to organizational security policy. Standardizing the tools reduces incident setup time and produces more comparable evidence across clusters. Those images should be versioned and patched like any other privileged operational artifact.

Finally, know when to escalate. If the failure is inside a managed control plane, cloud load balancer, storage backend, or other provider-owned layer, collect timestamps, object names, request IDs, and reproducible symptoms before opening the provider case. Good evidence shortens escalation and prevents repeated first-line troubleshooting.

Within the CNCF administration model, troubleshooting is a structured search through object state, events, logs, metrics, and component ownership. The strongest operators do not memorize a single command for every failure; they use the architecture to decide what evidence to collect next.

Filed under DevOps & Automation