INSIGHTS
Cybersecurity

CNCF KCSA: Runtime Security for Kubernetes Clusters

In this article
  1. Start with a runtime threat model
  2. Reduce the runtime attack surface before detecting it
  3. Use seccomp and kernel controls to restrict behavior
  4. Collect runtime telemetry that preserves context
  5. Define detections around deviations from expected behavior
  6. Monitor Kubernetes API behavior as part of runtime security
  7. Protect the node and container runtime interface
  8. Integrate runtime security with incident response
  9. Test detections with safe, known behaviors

Kubernetes security does not end when an image passes a registry scan and a Pod is admitted. Once a container is running, the important questions change: what processes actually execute, which files are opened, what network connections are made, which syscalls are attempted, and whether the workload behaves like the software that was reviewed. Runtime security is the discipline of observing and constraining that live behavior.

The current KCSA includes Kubernetes threat modeling, platform security, observability, and runtime-oriented controls. A practical runtime program combines prevention, telemetry, detection, and incident response rather than relying on a single agent or rule set.

Start with a runtime threat model

Runtime controls should be tied to plausible failure paths. An attacker may exploit a vulnerable web application, steal a service-account token, execute a shell inside a container, abuse a writable mount, escape through an unsafe kernel interface, or pivot to other workloads over the network. Each path leaves different evidence and is constrained by different controls.

Map trust boundaries between the container, node kernel, Kubernetes API, cloud metadata services, persistent storage, and adjacent services. If a workload is internet-facing and processes untrusted input, assume compromise is possible and focus on limiting what a compromised process can reach.

This mindset is more useful than treating runtime security as a list of signatures. The same shell command may be normal in an administrative toolbox image and highly suspicious in a minimal payment-service container.

Reduce the runtime attack surface before detecting it

The best runtime alert is often the action a workload cannot perform. Non-root execution, dropped Linux capabilities, read-only filesystems, seccomp, AppArmor or SELinux, restrictive mounts, minimal RBAC, and NetworkPolicy all remove options from an attacker after code execution.

Use small images and avoid shipping compilers, package managers, shells, or debugging utilities in production images unless the application truly needs them. Fewer binaries mean fewer post-exploitation tools and fewer packages to patch.

Container runtime choice also affects the operational surface. The differences discussed in container runtime behavior are a reminder that runtime implementation, isolation, logging, and node integration matter even when Kubernetes presents a common API.

Use seccomp and kernel controls to restrict behavior

Seccomp filters system calls a process can make into the Linux kernel. Kubernetes supports runtime-default and local seccomp profiles. For many workloads, a runtime-default profile is a strong baseline because it blocks syscalls that ordinary application processes rarely need.

Capabilities divide traditional root privileges into narrower units. Drop all capabilities by default and add back only what is justified. Pair that with allowPrivilegeEscalation: false where possible so a process cannot gain broader privilege through mechanisms such as setuid executables.

AppArmor and SELinux can add mandatory access restrictions beyond normal file permissions. These controls are especially valuable when an attacker obtains code execution but still needs to access host files, devices, or sensitive mounts.

Collect runtime telemetry that preserves context

Useful telemetry includes process execution, container identity, image digest, Kubernetes namespace and Pod metadata, network connections, file changes, authentication events, API activity, and node-level security events. A process name without workload context is difficult to investigate in an environment where Pods are constantly created and destroyed.

Correlate runtime observations with Kubernetes metadata. If an unexpected process appears, responders should be able to identify the Deployment, service account, node, image, owner, and recent rollout history quickly.

Retain enough telemetry to investigate delayed detections. Short-lived containers may disappear long before an analyst reviews an alert, so central collection is often necessary.

Define detections around deviations from expected behavior

Strong detections often describe behaviors that should be rare: a shell spawned inside an application container, writes to sensitive host paths, access to container-runtime sockets, package installation in production, unexpected outbound connections, creation of privileged Pods, or reads of service-account tokens by processes that normally never touch them.

Signature-based vulnerability knowledge still matters. Understanding CVEs and vulnerability identifiers helps teams connect exploit intelligence to images and runtime exposure. But a CVE alone does not prove exploitation, and behavior detections can catch unknown or misconfiguration-driven attacks.

Tune rules against workload baselines. If every deployment generates dozens of harmless alerts, responders will stop trusting the system.

Monitor Kubernetes API behavior as part of runtime security

A compromised container may use its service-account token to query Secrets, list Pods, create Jobs, or modify workloads. Those actions happen through the Kubernetes API and may not be visible in process-only telemetry. Audit logs and API server observations are therefore an important runtime signal.

Watch for unusual verbs, sensitive resource access, cross-namespace enumeration, token use from unexpected sources, and creation of privileged or host-mounted workloads. The strength of these detections depends on knowing what the service account normally does.

Minimal RBAC reduces both risk and noise. A workload that only has permission to read one ConfigMap cannot turn a stolen token into cluster-wide enumeration.

Protect the node and container runtime interface

Containers ultimately depend on the node kernel and runtime. Access to the container runtime socket, host PID namespace, host network, hostPath mounts, or privileged mode can collapse isolation. Restrict these features through admission policy and strong administrative boundaries.

Patch node operating systems, kernels, runtimes, and Kubernetes components. Node vulnerabilities can turn an application compromise into a broader escape or persistence opportunity. General patch-management discipline applies, but Kubernetes adds scheduling and availability concerns that require controlled draining and maintenance.

Separate routine workload administration from node-level access. The ability to deploy an application should not automatically grant SSH or privileged debug access to every worker.

Integrate runtime security with incident response

When a runtime alert fires, preserve evidence before deleting the Pod. Capture metadata, logs, process details, image digest, network connections, relevant audit events, and the owning controller. If the workload is ephemeral, deletion may destroy the only local evidence.

Containment can include scaling a deployment to zero, isolating a namespace with NetworkPolicy, revoking credentials, cordoning a node, or blocking an image digest. Choose the smallest action that stops harmful activity without unnecessarily destroying evidence or causing a wider outage.

The stages in cyber incident response map well to container platforms: prepare, detect, contain, eradicate, recover, and improve controls using what the incident taught you.

Test detections with safe, known behaviors

Do not wait for a real compromise to discover that telemetry is missing. In a controlled environment, run a shell in a container, access a forbidden path, attempt a disallowed network connection, or create a workload that should violate admission policy. Confirm that expected prevention and detection signals appear.

Test after upgrades to nodes, runtime agents, CNI plugins, and logging pipelines. Security tooling that silently loses metadata after a platform change can leave the team blind while dashboards still appear healthy.

Keep tests representative of production architecture. A detection proven only in a single-node lab may behave differently with managed nodes, service mesh, hardened kernels, or different runtime implementations.

Runtime security should feed earlier stages of the lifecycle. If a workload repeatedly needs a shell to perform operational tasks, improve the operational interface instead of permanently shipping a full toolbox. If runtime telemetry shows an application reaching an undocumented external service, update architecture and egress policy.

Patterns in incidents can become admission rules, image-hardening requirements, or CI checks. That closes the loop between detection and prevention.

Within CNCF certifications, runtime security demonstrates why cloud-native defense is layered. Image scanning, admission, process restrictions, network segmentation, API controls, telemetry, and response all cover different stages of the attack path.

Runtime baselines should be workload-specific. A Java service, an NGINX proxy, and a database have different normal process trees and network behavior. Create expectations from known-good deployments and version them with the application so detections can evolve when releases legitimately change behavior.

Process execution is especially useful in immutable-image environments. If the container image normally launches one application binary, later execution of bash, curl, nc, a package manager, or a compiler deserves attention. The same process may be normal in a toolbox namespace, which is why namespace and image context must accompany the alert.

File-integrity monitoring can focus on paths that should not change after startup. Writes under application binaries, startup scripts, SSH configuration, or credential directories may indicate persistence or tampering. Avoid monitoring every temporary file; high-volume noise hides the changes that matter.

Network detections should distinguish service communication from discovery behavior. A workload that suddenly scans many cluster IPs, connects to the API server despite having no management role, or reaches an unfamiliar external ASN may warrant investigation. Pair network telemetry with NetworkPolicy so unexpected paths are both visible and, where practical, blocked.

Cloud metadata endpoints require special attention. In some environments, a compromised Pod can attempt to reach instance metadata and obtain node or workload credentials. Use provider-specific protections, network controls, and workload identity so application Pods do not inherit more cloud authority than necessary.

Service-account token handling has changed over Kubernetes history, and modern clusters can use projected, time-limited tokens. Prefer short-lived credentials and disable automatic token mounting for workloads that do not call the Kubernetes API. Runtime detection should flag unexpected reads of token files or API calls from identities that normally have no such behavior.

Privileged debugging is another runtime risk. Ephemeral containers and node debug features are extremely useful for incident response, but they should be limited to trusted operators and audited. A malicious actor with debug privileges may access process namespaces, filesystems, or host resources that ordinary workloads cannot reach.

Node-level agents need their own threat model because they often run with host mounts, privileged access, or broad visibility. Protect their images and update paths, restrict who can modify DaemonSets, and alert on unexpected changes. Compromise of a cluster-wide security agent can create both visibility gaps and a powerful persistence mechanism.

Kernel signals such as denied syscalls, namespace changes, capability use, and executable mappings can be high value, but collection overhead must be understood. Measure CPU, memory, event volume, and storage impact of security sensors under real workload load. A control that destabilizes nodes will eventually be disabled.

Detection pipelines should survive cluster incidents. If all telemetry is stored inside the same cluster, an attacker or outage that affects the cluster may also erase evidence. Forward important logs and alerts to a separately controlled destination with retention appropriate to investigation requirements.

Incident playbooks should define when to isolate a Pod, namespace, or entire node. A single suspicious shell in one stateless web Pod may justify deleting and replacing that workload after evidence capture, while signs of container escape or node compromise should trigger a much broader response. Scope containment according to the trust boundary that may have failed.

After containment, rebuild from trusted artifacts rather than attempting to “clean” a compromised container. Containers are designed to be replaceable. If the node itself may be compromised, rebuilding or reimaging the node is generally more defensible than trusting ad-hoc cleanup.

Runtime security programs improve when developers can see why controls fired. Provide enough context in alerts to identify the offending process, image, namespace, deployment, and rule rationale. Actionable feedback turns security from an opaque gate into a source of engineering improvement.

Runtime detections should include container lifecycle context. A process that appears for a few seconds during an init container may be expected, while the same process in the long-running application container may be suspicious. Preserve container name and lifecycle phase in telemetry.

Monitor use of Kubernetes exec, port-forward, attach, and ephemeral-container features because they provide legitimate operators with deep workload access. Unexpected use by a rarely seen identity or against sensitive namespaces can be an important signal even when no exploit is involved.

Fileless behavior still matters. An attacker can execute code in memory or use existing binaries without writing a new executable. This is why process trees, network behavior, API activity, and syscall observations complement file-integrity monitoring rather than depending on it.

Container restarts can erase volatile evidence. If a suspicious Pod is repeatedly crashing, capture previous logs and runtime events before the controller replaces it. For high-value workloads, centralize logs with timestamps precise enough to correlate process, API, and network events.

Threat hunting can use hypotheses rather than alerts alone. Search for Pods that recently contacted the API server, containers that launched interactive shells, unusual outbound DNS patterns, or nodes that ran privileged debug workloads. Periodic hunts reveal visibility gaps and weak assumptions before an incident forces the issue.

Measure response quality. Track time from runtime alert to workload identification, containment, credential rotation, trusted rebuild, and restoration. Improvements in these intervals show whether the runtime program is becoming operationally useful rather than merely generating more events.

Runtime programs also need ownership for false positives and policy drift. Every high-severity rule should have a team that can explain the signal, tune legitimate behavior, and maintain it as Kubernetes versions and workloads change. Rules without owners usually become noisy, disabled, or ignored.

Use a small set of high-confidence detections before chasing full coverage. Reliable alerts for privileged container creation, unexpected shell execution, sensitive host mounts, suspicious API access, and known-bad outbound behavior can provide more operational value than hundreds of poorly tuned rules.

Keep the response path tested end to end. Security teams should know who can isolate a workload, revoke its credentials, preserve node evidence, and authorize a rebuild, so a high-confidence runtime alert leads to coordinated action instead of a long search for ownership.

Filed under Cybersecurity