INSIGHTS
DevOps & Automation

CNCF CKA: Kubernetes Cluster Architecture

In this article
  1. Think of the control plane as the cluster's coordination layer
  2. Use the API server as the front door to cluster state
  3. Protect etcd because it stores authoritative cluster data
  4. Let the scheduler place unscheduled Pods on suitable nodes
  5. Use controllers to reconcile actual state with desired state
  6. Understand what runs on worker nodes
  7. Design high availability around failure domains, not just replica counts
  8. Keep add-ons separate from Kubernetes core components
  9. Plan upgrades around component compatibility and API evolution

Kubernetes architecture is built around a control plane that stores desired state and a set of worker nodes that run application Pods. That simple sentence is the foundation for most operational reasoning in Kubernetes. Administrators declare what should exist through the API, controllers reconcile the cluster toward that state, the scheduler decides where workloads belong, and node components make those workloads run.

The current CKA gives 25% of the exam to cluster architecture, installation, and configuration. That weighting reflects how central the architecture is: troubleshooting networking, scheduling, storage, or workloads is easier when you know which component owns each decision.

Think of the control plane as the cluster’s coordination layer

The control plane manages cluster state and decisions. Its major components include the API server, etcd, the scheduler, and controller managers. In production, these components are usually deployed redundantly so that one machine failure does not remove the cluster’s management plane.

The control plane does not directly execute most application containers. Instead, it records desired state and coordinates the nodes that run workloads. This separation lets Kubernetes manage applications across many machines while maintaining one consistent API.

Managed Kubernetes services may hide the machines that host these components, but the logical responsibilities remain the same. Understanding them still matters when deciding whether a problem belongs to the application, the cluster, or the managed-service provider.

Use the API server as the front door to cluster state

The kube-apiserver exposes the Kubernetes API. kubectl, controllers, operators, schedulers, nodes, and automation communicate through that API rather than modifying etcd directly. The API server handles request processing, authentication, authorization, admission, validation, and persistence of accepted objects.

This makes API availability critical for management operations. Existing Pods may continue serving traffic during a temporary API outage, but new scheduling, scaling, configuration changes, and controller reconciliation can be disrupted.

Security also converges on the API server. Certificates, authentication mechanisms, RBAC, and admission policy determine who can ask the cluster to do what. A cluster with weak API authorization has a weak control plane regardless of how secure individual applications appear.

Protect etcd because it stores authoritative cluster data

Kubernetes stores its authoritative cluster state in etcd, a strongly consistent key-value database. The API server persists objects there, including configuration, workload specifications, and other state required to reconstruct what the cluster is supposed to look like.

That makes etcd backup and recovery a fundamental administrative responsibility in self-managed clusters. A failed application Pod can often be recreated by a controller, but a lost or corrupted cluster datastore can become a full control-plane recovery event.

Access to etcd should be tightly restricted and protected with TLS. Administrators should monitor latency, free space, quorum health, and backup success. Backups need restore testing; a snapshot that has never been restored is only an assumption about recoverability.

Let the scheduler place unscheduled Pods on suitable nodes

The kube-scheduler watches for Pods that do not yet have a node assignment and evaluates candidate nodes. It considers resource requests, taints and tolerations, node selectors, affinity rules, topology constraints, and other scheduling policies before binding a Pod to a node.

The scheduler does not start containers. It makes a placement decision. The kubelet on the selected node is responsible for making the assigned Pod run. This distinction helps with Pending Pods: a scheduling failure happens before container startup, while an image or runtime failure happens after placement.

Administrators should learn to read scheduling events rather than guessing. Messages about insufficient CPU, unmatched affinity, untolerated taints, or volume binding point directly to the class of constraint blocking placement.

Use controllers to reconcile actual state with desired state

Controllers are continuous control loops. A Deployment controller observes a Deployment and manages ReplicaSets. A ReplicaSet controller ensures the required number of Pods exists. Node and endpoint controllers maintain other parts of cluster state. If reality drifts from the declaration, controllers take action to converge again.

This reconciliation model explains Kubernetes self-healing. The system is not running one long orchestration script that fails permanently if step seven breaks. Controllers repeatedly observe state and retry from the current situation.

Operators and custom controllers extend the same model. They watch Kubernetes objects and implement domain-specific behavior. That power makes extension easy, but a poorly designed controller can also create excessive API load or repeatedly fight another controller over the same resource.

Understand what runs on worker nodes

Worker nodes host application Pods and the agents required to integrate the machine with Kubernetes. The kubelet watches Pod specifications assigned to the node and works with the container runtime to create containers, mount volumes, configure required resources, and report status back to the API.

Container runtimes implement the Container Runtime Interface and handle the lifecycle of containers and images. Networking components establish Pod networking, while a Service proxy implementation such as kube-proxy or an alternative data plane makes Service behavior available.

The relationship between Pods and containers is explained more deeply in Pods versus containers: the Pod is the Kubernetes scheduling and workload unit, while the runtime actually executes the containers inside it.

Design high availability around failure domains, not just replica counts

Running multiple control-plane instances improves resilience only when those instances are distributed intelligently. Three API servers on one failed host are not highly available. etcd quorum members should span appropriate failure domains, and load balancers must have healthy paths to API server replicas.

Worker availability matters too. Multiple replicas of an application provide little protection if the scheduler places all of them on one node or one failure domain. Topology spread and anti-affinity can help distribute workloads when the application requires stronger resilience.

High availability and disaster recovery are different. Redundant live components protect against individual failures, while backups and documented recovery procedures protect against corruption, operator error, or broader loss of state.

Keep add-ons separate from Kubernetes core components

Clusters often include DNS, metrics, ingress controllers, CSI storage drivers, CNI networking plugins, policy engines, and observability agents. These are essential in many real deployments, but they are not all the same thing as the core Kubernetes control plane.

This distinction matters during incidents. If CoreDNS is unhealthy, service-name resolution can fail while the API server remains healthy. If the CNI plugin fails on one node, new Pods may lack network connectivity even though scheduling succeeds. If a CSI driver is broken, Pods can remain Pending waiting for storage.

Architecture diagrams should therefore include both core components and critical add-ons so operators can trace dependencies rather than assuming “Kubernetes” is one monolithic service.

Plan upgrades around component compatibility and API evolution

Kubernetes releases evolve quickly. Control-plane components, kubelets, kubectl, admission webhooks, CNI plugins, CSI drivers, and custom controllers need compatible versions. Upgrade runbooks should follow supported version-skew policies and verify third-party components before changing the cluster.

API deprecations are especially important. A manifest or controller using an API version removed by the target release can fail even if the underlying application image has not changed. Scan workloads and extensions for deprecated APIs before an upgrade window.

Managed services automate some control-plane upgrades, but application and extension compatibility still belongs to the cluster user. Provider automation reduces operational work; it does not eliminate version awareness.

If kubectl cannot reach the cluster but existing applications are serving traffic, investigate API connectivity and control-plane availability. If a Pod is Pending, examine scheduling and storage events. If a scheduled Pod cannot start, inspect kubelet, runtime, image, volume, and configuration issues. If a Service has no traffic, inspect selectors, EndpointSlices, and the network data plane.

Those branches are easier to remember when each component has a clear responsibility. Random restarts destroy evidence and can add new symptoms. Architecture gives incidents a search order.

Control-plane communication follows a hub-and-spoke API pattern. Nodes and controllers communicate with the API server over authenticated HTTPS rather than exposing every control-plane component directly to the cluster network. This centralizes policy and simplifies the security model, but it also makes API connectivity a critical dependency for kubelets and controllers.

Kubelet is not merely a container launcher. It continuously reconciles the Pods assigned to its node, reports status, runs probes, manages mounted volumes with plugins, and coordinates with the runtime and network components. A healthy container runtime with a broken kubelet still produces a Kubernetes node problem because the control plane loses its node agent.

Static Pods are another architectural concept administrators should recognize. The kubelet can manage Pods defined from local manifests without a normal controller creating them through the API. Some cluster bootstrapping tools use static Pods for control-plane components. Mirror Pods may appear in the API for visibility even though the kubelet owns the actual lifecycle from the local manifest.

Certificates and component identities form part of cluster architecture. Kubelets, administrators, controllers, and API servers authenticate with credentials that need controlled issuance and rotation. Expired or incorrect certificates can create symptoms that resemble network failure, so certificate health should be part of control-plane maintenance.

Admission control evaluates authenticated and authorized API requests before approved objects are written to cluster state. Mutating admission can modify requests, while validating admission can reject workloads that violate policy. This is where organizations commonly enforce image, security, resource, or governance requirements that are broader than one application team’s manifest.

Cluster architecture should also account for failure during bootstrap. A new node needs network reachability, credentials, runtime configuration, and the correct cluster information before it can join. Automation should make those prerequisites reproducible; snowflake node configuration makes replacement slow and undermines the disposable-node model.

API request load can become an architecture concern at scale. Controllers that list resources too frequently, monitoring systems with aggressive queries, or poorly designed operators can overload the API server and etcd. Efficient watches, rate limits, and controller design protect the control plane from its own extensions.

Finally, architecture documentation should identify ownership. Managed platforms divide responsibility between provider and customer; self-managed clusters place more components under the platform team. Incident response is faster when teams know who owns API availability, etcd recovery, node images, CNI, CSI, DNS, ingress, and certificate rotation.

Cluster DNS is often deployed as an add-on but is so central to application behavior that it should be treated as critical infrastructure. CoreDNS or an alternative resolver watches Kubernetes state and answers service-discovery queries. If DNS is unhealthy, applications can fail broadly even when control-plane APIs and Pod networking are otherwise functional.

Cluster networking similarly spans core expectations and implementation-specific components. Kubernetes assumes Pod-to-Pod connectivity according to its network model, while the CNI plugin supplies the actual routes, overlays, or eBPF programs. Node replacement, upgrade, and troubleshooting procedures must therefore include the CNI version and configuration as first-class dependencies.

Time synchronization is another quiet dependency. Certificates, event timelines, distributed logs, and some authentication systems assume reasonably accurate clocks. Large clock drift can turn a healthy cluster into a collection of confusing TLS and ordering failures. Production architecture should include reliable time sources and monitoring for drift.

Resource isolation on nodes is part of architecture too. System daemons, kubelet, the container runtime, networking agents, and workloads all compete for CPU, memory, disk, and I/O. Node allocatable settings and reserved resources help prevent application Pods from consuming capacity required for cluster operation.

Architecture becomes operationally complete only when teams can rebuild it. Infrastructure-as-code, node images, bootstrap configuration, cluster manifests, and backup procedures should allow a replacement cluster or control plane to be created without relying on undocumented manual steps.

Node lifecycle is part of architecture, not merely fleet management. Nodes should be replaceable through automated provisioning, joining, draining, and removal. Long-lived manual nodes accumulate configuration drift and make upgrades risky because nobody knows which local change is required for workloads to survive.

Cluster capacity planning should reserve headroom for failures and maintenance. If every node is packed to its requested capacity, draining one node for an upgrade may leave its Pods unschedulable. Architecture should include enough spare capacity or elastic scaling to absorb routine disruption.

Control-plane recovery procedures should be practiced in an isolated environment. Verify that etcd snapshots, certificates, manifests, and load-balancer configuration are sufficient to restore API access. Recovery documentation should specify version compatibility and the order in which components return. A backup is valuable only when operators can convert it back into a working control plane under realistic constraints.

Architecture reviews should include dependencies outside the cluster as well. Image registries, identity providers, cloud APIs, DNS, time services, and external load balancers can all prevent a healthy Kubernetes control plane from delivering a healthy application platform. A cluster diagram that ends at the node boundary is incomplete for production operations.

Capacity and dependency ownership should be documented before upgrades or incidents begin.

Operational readiness also means knowing the cluster’s single points of dependency even when components are replicated. Shared network paths, one certificate authority, one external identity provider, or one cloud-region API can still become a common failure. High availability should be evaluated by dependency graph, not only by counting replicas.

That clarity turns component knowledge into repeatable operational decisions during both change and failure.

For learners moving from KCNA foundations toward administration, the CNCF model is consistent: declare desired state through the API, let controllers and the scheduler coordinate it, and let node components execute the workload. That mental model supports almost every deeper Kubernetes topic.

Filed under DevOps & Automation