INSIGHTS
Networking

CNCF CKA: Kubernetes Services and Networking

In this article
  1. Start with the Kubernetes network model
  2. Use Services to decouple clients from Pod churn
  3. Know what each Service type changes
  4. Understand selectors and EndpointSlices before blaming the network
  5. Use cluster DNS for service discovery
  6. Understand how Service traffic reaches backends
  7. Separate east-west networking from north-south access
  8. Troubleshoot from application endpoint outward
  9. Operate networking as a system, not as isolated YAML

Kubernetes networking lets Pods communicate across nodes while Services provide stable endpoints over a changing set of Pods. The key operational challenge is that application replicas are replaceable: Pod IP addresses can disappear during a rollout, scaling event, or node failure. Clients therefore need a stable way to find the application without tracking every backend directly.

The current CKA assigns 20% of the exam to Services and networking. Administrators need more than definitions of ClusterIP and LoadBalancer. They need to understand selectors, EndpointSlices, DNS, the cluster network, external access, and how to troubleshoot traffic from client to Service to Pod.

Start with the Kubernetes network model

Each Pod receives its own cluster-wide IP address from the cluster network. Containers in the same Pod share the Pod network namespace and can communicate over localhost, while different Pods communicate using their Pod IP addresses when network policy permits it.

The cluster networking implementation is provided by the chosen CNI plugin and underlying infrastructure. Kubernetes defines the model but does not prescribe one universal packet-processing technology. Some clusters use overlay networking, some integrate closely with cloud routing, and others use eBPF-based data planes.

Operators should know which implementation their cluster uses because packet paths, troubleshooting commands, NetworkPolicy support, and observability can differ materially across platforms.

Use Services to decouple clients from Pod churn

A Service represents a stable network endpoint for one or more backend Pods. Most Services use a label selector to identify matching Pods. Kubernetes tracks those backends through EndpointSlice objects so the data plane knows where to forward traffic.

This abstraction allows Deployments to replace Pods without forcing every client to rediscover individual addresses. As long as new replicas become ready and match the selector, the Service continues to represent the application.

That stability is one reason the Pod-versus-container model matters. The Pod is disposable from a client perspective; the Service is the long-lived access point.

Know what each Service type changes

ClusterIP is the default Service type and exposes the Service only inside the cluster network. It is the normal choice for internal application components. NodePort allocates a port on cluster nodes and forwards traffic to the Service, which can be useful as a building block but is rarely the most elegant public-access model by itself.

LoadBalancer asks an integrated platform to provision an external load balancer. The resulting implementation varies by cloud or environment. The general principles of load balancing still apply: clients need a stable address, health must reflect usable backends, and failure domains matter.

ExternalName is different: it provides a DNS alias rather than proxying traffic to Kubernetes endpoints. Use it carefully because application protocols and TLS certificates still see the external hostname behavior.

Understand selectors and EndpointSlices before blaming the network

A Service with the wrong selector may exist perfectly while routing to zero Pods. When traffic fails, inspect the Service selector and compare it with Pod labels. Then inspect EndpointSlices to confirm that ready backends are registered.

Readiness affects this path. A running Pod that is not ready may be excluded from normal Service traffic. That is desirable when the application has started but is not yet safe to receive requests.

Services can also exist without selectors, with endpoints managed manually or by another controller. Those designs are useful for external systems or migration scenarios but remove some of the automatic relationship between Pod labels and backend registration.

Use cluster DNS for service discovery

Kubernetes normally provides DNS names for Services and Pods through a cluster DNS add-on. Applications can refer to a Service by name instead of hard-coding its virtual IP. Namespace-aware names make service discovery predictable across environments.

DNS introduces its own failure domain. If CoreDNS or equivalent components are unhealthy, applications may report connection failures even when the Service and Pods are otherwise healthy. Test name resolution separately from raw IP connectivity during incidents.

The same fundamentals behind DNS caching still matter. Stale records, resolver behavior, and negative caching can make application behavior appear inconsistent during changes.

Understand how Service traffic reaches backends

A Service virtual IP is not normally an application process listening on that address. A Service proxy implementation watches Service and EndpointSlice state and programs packet-processing rules that steer traffic toward backends. Historically kube-proxy commonly used iptables or IPVS, while some modern CNI implementations provide their own service proxying.

This detail matters during troubleshooting because a Service can look correct in the API while node-level packet processing is broken. One node may have stale rules or a malfunctioning network agent, producing failures only for clients or backends that traverse that node.

Administrators should know the cluster’s chosen data plane before reaching for low-level commands. Troubleshooting iptables on a cluster that uses an eBPF proxy may waste time and lead to the wrong conclusions.

A headless Service sets the cluster IP to None. Instead of providing one virtual IP, DNS can return records associated with the individual endpoints. Stateful systems may use this pattern when clients need stable identities or direct awareness of replicas.

Headless Services do not mean “no Kubernetes networking.” They still provide service discovery and can still be paired with selectors and EndpointSlices. The difference is that clients see backend addresses more directly rather than sending traffic to a single virtual Service IP.

This model is common with StatefulSets, databases, and clustered applications where each replica has an identity that matters to the protocol.

Separate east-west networking from north-south access

Services primarily solve stable access inside or around the cluster. External HTTP routing is usually handled by Ingress or Gateway API, while a LoadBalancer Service can expose a Service directly. NetworkPolicy controls which flows are permitted among Pods and external peers when the CNI implementation supports it.

These features are complementary. A Gateway can correctly route traffic to a Service while NetworkPolicy blocks the resulting Pod connection. A Service can have healthy endpoints while a public firewall prevents clients from reaching the load balancer.

Draw the traffic path from client to backend and map each policy point rather than treating “networking” as one undifferentiated layer.

Troubleshoot from application endpoint outward

Start with the backend Pods. Are they ready and listening on the expected container port? Next check the Service selector, ports, targetPort, and EndpointSlices. Test the Service from another Pod. Then test DNS by name and compare that result with direct Service IP access.

If internal Service access works, move outward to NodePort, LoadBalancer, Ingress, or Gateway layers. Check firewall and security-group rules, external addresses, health probes, and route status. This inside-out method isolates the failing layer faster than changing multiple resources at once.

For application-focused learners, CKAD workload concepts provide useful complementary context because Services, readiness, and application ports are tightly connected.

Operate networking as a system, not as isolated YAML

Stable Kubernetes networking depends on several cooperating components: the CNI plugin, node networking, Services, EndpointSlices, DNS, proxies or eBPF data planes, external load balancers, and policy. A manifest can be syntactically correct while one of those dependencies is unhealthy.

Monitor DNS latency, packet drops, connection failures, node network health, and load-balancer backend status. Record which component owns public IPs, DNS, certificates, and firewall policy so incident responders know where to look.

Service ports deserve careful reading because three port numbers may appear in one path. The Service port is the port clients use against the Service. targetPort identifies the port on the backend Pod, either numerically or by a named container port. NodePort adds another externally reachable node-level port when that Service type is used. Misaligning these values is a common cause of silent-looking connectivity failure.

EndpointSlice readiness reflects backend usability, but special cases exist for terminating or serving endpoints. Modern controllers can retain richer endpoint condition information than the older Endpoints object. Operators should prefer EndpointSlice-aware tooling when diagnosing large or modern clusters.

Service session affinity can keep a client mapped to the same backend for a period when ClientIP affinity is configured. This can help some legacy applications, but it is not a substitute for proper shared state or application-level session design. Sticky traffic can also produce uneven load when a few clients generate most requests.

externalTrafficPolicy influences how external Service traffic is handled and whether original client source addresses are preserved in some environments. Local policy can avoid an extra cross-node hop and preserve source information, but nodes without local healthy endpoints may behave differently. The right setting depends on the load balancer and application requirements.

NetworkPolicy should be tested rather than assumed. Kubernetes exposes the API, but actual enforcement depends on the cluster networking implementation. A policy object can exist while having no effect if the chosen network plugin does not implement the feature. Platform documentation should state which policy capabilities are supported.

Dual-stack clusters add IPv4 and IPv6 considerations. Services and Pods may have addresses from one or both families depending on configuration. DNS can return A and AAAA records, and external systems must be ready for the chosen address families. IPv6 troubleshooting should not assume that familiar IPv4 NAT behavior applies in the same way.

MTU mismatches can cause confusing partial connectivity, especially with overlays, tunnels, or service meshes that add headers. Small requests may succeed while larger packets stall or fragment. When connectivity problems depend on payload size, inspect path MTU and encapsulation overhead rather than focusing only on Service selectors.

Network observability should include packet drops, DNS errors, connection latency, load-balancer health, and node-level interface state. Application metrics alone may show elevated errors without explaining whether the failure occurred before the request reached the Pod. A layered dashboard is more useful than one generic “network healthy” indicator.

Services use labels for backend selection, so label governance affects networking reliability. Reusing generic labels such as app=web across unrelated components can accidentally place the wrong Pods behind a Service. Use a consistent labeling scheme that distinguishes application, component, instance, and environment where needed.

Readiness gates can extend ordinary readiness by requiring additional conditions before a Pod is considered ready. Integrations such as external load-balancer controllers may use this to avoid sending traffic before cloud infrastructure has completed registration. When readiness seems stuck, inspect both container probes and any custom readiness conditions.

Connection tracking can influence rollout behavior. Long-lived TCP connections may continue using an old backend even after new Pods become ready, depending on protocol and proxy behavior. Applications that require rapid traffic migration should test connection lifetime, graceful termination, and load-balancer draining explicitly.

Service troubleshooting should include NetworkPolicy and host firewalls. A correct Service path can still be blocked at the Pod, node, subnet, or cloud-security layer. Keep a documented list of policy systems in the traffic path so responders know which controls can affect the same connection.

Multi-cluster service discovery adds another layer and should be adopted only when required. Cross-cluster gateways, global DNS, service exports, and meshes can provide regional resilience, but they also introduce identity, latency, and failover complexity. A reliable single-cluster Service model is the foundation before extending discovery across clusters.

Health probes used by an external load balancer may not follow the same path as real user traffic. A probe can succeed against one node or port while the application route still fails. Validate representative client traffic in addition to infrastructure health checks, especially after changes to source-IP policy or gateway topology.

Application protocols can also influence Service troubleshooting. HTTP errors prove that a request reached an HTTP-speaking endpoint, while a TCP timeout suggests a lower-layer connectivity problem. TLS handshake errors sit between those cases and may indicate certificates, server names, or interception. Classifying the failure by protocol stage narrows the search quickly.

Keep network changes reversible. CNI upgrades, Service proxy changes, or cluster-wide policy can affect every workload. Stage changes on a subset of nodes or clusters when possible and preserve a known rollback path rather than testing a new data plane for the first time in production.

Document address ranges and routing ownership. Pod CIDRs, Service CIDRs, node networks, VPNs, and on-premises networks must not overlap unexpectedly. Address overlap can create intermittent or one-way connectivity that is difficult to diagnose from Kubernetes objects alone. Network planning before cluster creation is far easier than renumbering live clusters after routes collide.

Service design should also consider failure during rolling updates. If readiness becomes true too early, new Pods may receive traffic before caches, routes, or dependencies are ready. If termination removes endpoints too late, clients may continue sending requests to a shutting-down process. Networking reliability therefore depends on probe and lifecycle design as much as on packet routing.

For production changes, capture a baseline before modifying CNI, DNS, kube-proxy, or Service configuration. Compare latency, error rates, endpoint counts, and DNS resolution after the change. A measured before-and-after view is much more reliable than deciding that networking “looks fine” because Pods are still Running.

Reliable networking is ultimately an operational contract between labels, endpoints, data-plane components, and the applications that depend on them.

Clear traffic-path ownership also makes planned changes safer and incident handoffs much faster.

Within the CNCF ecosystem, the durable model is straightforward: Pods have routable identities, Services provide stable discovery over changing backends, and additional APIs expose or restrict that traffic. Understanding those layers is more valuable than memorizing a single vendor’s network implementation.

Filed under Networking