INSIGHTS
Networking

CNCF KCSA: Network Policies for Kubernetes Workloads

In this article
  1. Verify that your networking implementation actually enforces policy
  2. Understand how Pod selection creates isolation
  3. Use default-deny as a starting state, not an end state
  4. Write ingress rules around the service relationship
  5. Use egress policy to contain compromised workloads
  6. Treat DNS as an explicit dependency
  7. Test policy from both allowed and denied paths
  8. Design namespace labels as security boundaries
  9. Know the limits of NetworkPolicy

Kubernetes makes it easy for applications to discover and reach one another, but default connectivity is often broader than a security-conscious production design needs. NetworkPolicy provides a native API for declaring which Pod traffic should be allowed at layer 3 and layer 4. Used well, it changes the cluster from an environment where compromise can spread freely into one where network paths are deliberately constrained.

The current KCSA includes isolation, segmentation, NetworkPolicy, Kubernetes trust boundaries, and attacker-on-the-network scenarios. The difficult part is not writing one YAML object. It is understanding selection, ingress and egress isolation, namespace relationships, DNS dependencies, CNI enforcement, and how to test policy without breaking legitimate service communication.

Verify that your networking implementation actually enforces policy

NetworkPolicy is a Kubernetes API, but enforcement is performed by the cluster networking implementation. A cluster can accept NetworkPolicy objects while using a network plugin that does not implement them, producing a dangerous illusion of protection. Confirm support in the CNI or managed-service documentation and test behavior with controlled traffic.

Policy behavior also depends on implementation details and supported features. Standard Kubernetes NetworkPolicy handles IP- and port-level rules, not arbitrary application-layer authorization. If you need HTTP paths, user identity, or service-to-service authentication, those controls belong elsewhere in the architecture.

Start by understanding the current flat network. Inventory namespaces, application tiers, shared services, external dependencies, monitoring collectors, admission services, and DNS. A useful policy model comes from real communication requirements rather than from a generic template.

Understand how Pod selection creates isolation

A NetworkPolicy applies to Pods selected by its podSelector in the policy’s namespace. Merely creating a policy does not isolate every Pod. A Pod becomes isolated for ingress or egress when at least one policy selects it for that direction. Once isolated, allowed traffic is the union of the rules from all policies that select the Pod.

This additive model matters. Two policies do not override each other in sequence like traditional firewall rules. If either selected policy permits a path, that path is allowed. Troubleshooting therefore requires looking at all policies that select the Pod, not only the one with the most obvious name.

Labels are security inputs. If access is granted to Pods with role=frontend, who can change that label? Admission policy, RBAC, and deployment ownership determine whether the selector can be manipulated. NetworkPolicy is strongest when label governance is also controlled.

Use default-deny as a starting state, not an end state

A common pattern is to establish default-deny ingress, egress, or both, then add narrowly scoped allow policies. This is easier to reason about than trying to enumerate every dangerous destination while leaving all other traffic open.

Do not roll out default-deny blindly. Applications may rely on DNS, time synchronization, identity providers, metrics endpoints, message brokers, package mirrors, cloud metadata services, or external APIs. If those dependencies are not mapped, the policy can convert a security improvement into a production outage.

Begin in a test namespace or representative environment. Capture existing flows, document intended dependencies, apply default-deny, and add permits one service relationship at a time. The general logic of network segmentation carries over: boundaries create value only when traffic between segments is explicitly understood.

Write ingress rules around the service relationship

Ingress policy should answer a concrete question: which sources need to reach this workload, on which protocol and port? A backend API might accept traffic from frontend Pods in the same application namespace, from an ingress controller namespace, and from an observability collector. Those are different trust relationships and should be visible in policy.

podSelector can select source Pods within a namespace. namespaceSelector can select namespaces based on labels, and the two can be combined to express “Pods with this label in namespaces with that label.” Be precise about YAML structure because separate peer entries can mean a broader union than one combined peer.

IP blocks can represent external CIDRs, but their behavior can be affected by address translation and how the networking implementation sees source IPs. Prefer identity through Pod and namespace selectors for in-cluster workloads when possible.

Use egress policy to contain compromised workloads

Ingress filtering receives much attention, but egress controls can be equally important. If an internet-facing application is compromised, unrestricted egress may let an attacker reach internal databases, the cloud metadata endpoint, external command-and-control infrastructure, or arbitrary data-exfiltration destinations.

Design egress around required destinations. Permit DNS to the cluster resolver, application calls to known services, and specific external ranges or ports where the architecture supports stable addressing. Be cautious with broad internet CIDRs because they can undermine the containment goal.

Some external services use changing IP ranges or CDNs, making pure layer-3 allowlisting difficult. In those cases, NetworkPolicy may need to be combined with egress gateways, proxies, service-mesh controls, or provider-native firewall mechanisms.

Treat DNS as an explicit dependency

One of the most common failures after enabling egress isolation is broken name resolution. A Pod may be allowed to contact its database IP but cannot resolve the database hostname because UDP or TCP DNS traffic to CoreDNS is blocked. The result looks like an application or service outage rather than a network-policy error.

Identify the cluster DNS Pods or Service and permit the protocols and ports actually required. If your environment uses node-local DNS caching or another resolver path, account for that architecture rather than copying a CoreDNS-specific example.

After changing policy, test both name resolution and the application connection. Successful dig or nslookup does not prove the destination port is reachable, and successful IP connectivity does not prove production DNS will work.

Test policy from both allowed and denied paths

Security testing should prove what is blocked, not only what still works. Create controlled source Pods with labels that should be allowed and denied, then attempt the same connection. Test across namespace boundaries, after label changes, and against both Service and direct Pod endpoints when appropriate.

Use packet capture, CNI observability, or flow logs if your platform provides them. NetworkPolicy failures can resemble application timeouts, so visibility into drops shortens troubleshooting. The principles in network monitoring logs apply even when the enforcement point is inside the cluster.

Include policy tests in deployment pipelines for high-value services. A regression that accidentally broadens a selector can be caught before it becomes the new production baseline.

Design namespace labels as security boundaries

Namespace selectors are powerful because they let policies describe trust zones such as production, shared infrastructure, or monitoring. That power means namespace labels must be governed. If ordinary application developers can add the label that grants access to a payment database, the policy boundary is weaker than it appears.

Use RBAC and admission controls to restrict modification of security-significant labels. Prefer labels whose meaning is clear and stable, such as an environment or platform-managed trust classification, rather than labels that reflect temporary team names.

Document who owns shared namespaces. Observability, ingress, service mesh, and security agents often need broad communication. Their access should be explicit rather than hidden inside a catch-all rule that every namespace eventually copies.

Know the limits of NetworkPolicy

NetworkPolicy does not authenticate an application user, validate TLS certificates, scan payloads, or understand SQL permissions. It controls network reachability. A database still needs authentication and authorization even when only one namespace can reach its port.

It also does not automatically protect host-networked processes or every path that bypasses normal Pod networking. Node firewalls, cloud security groups, service-mesh policy, and application controls may be part of a layered design.

Do not mistake network isolation for endpoint security. A permitted frontend can still attack a backend if the frontend is compromised and the backend trusts every request. Mutual authentication and least-privilege application identities remain important.

Network requirements change with applications. A new metrics collector, queue, or third-party API can make an old policy incomplete. Store policies with the application or platform configuration, review them alongside code changes, and make ownership clear.

During incidents, compare desired policy with the policies actually selecting a Pod. Watch for stale labels, old policy objects, and namespace moves. Because policies are additive, an overlooked legacy allow policy can leave a path open even after a stricter new policy is deployed.

Within CNCF certifications, NetworkPolicy is a good example of cloud-native security thinking: start from the communication graph, isolate by default, grant explicit paths, validate enforcement, and monitor the result. The YAML is only the expression of that reasoning.

Well-designed NetworkPolicy can make incident response faster because teams already know which paths should exist. If a compromised workload attempts unexpected east-west traffic, flow telemetry can reveal the deviation and policy can block many destinations automatically.

Emergency isolation should still be planned. Operators may need a rapid way to deny egress from a namespace, block a compromised tier, or permit a temporary forensic collector. Define those procedures before an incident so responders do not improvise broad policy changes under pressure.

After containment, review whether the attempted path was blocked because of deliberate policy or merely because the destination happened to be unavailable. Security maturity comes from controls whose expected behavior has been tested, documented, and observed under realistic failure conditions.

Policy design should account for health checks and platform traffic. Kubelet probes, load-balancer health checks, node-local components, and service-mesh sidecars may create flows that are not obvious from the application diagram. Test the actual path used by readiness and liveness checks before enforcing deny rules broadly.

Stateful services often have peer-to-peer communication requirements in addition to client traffic. Databases, message brokers, and distributed caches may need replication, leader election, or gossip between members. A policy that only allows application clients can leave the service partially healthy while internal clustering fails.

Use namespace and Pod selectors to express application intent, but verify how selectors behave when labels are absent. A typo in a namespace label can make an expected peer disappear from the allowed set. Admission checks for required labels can prevent this class of security and availability error.

External egress is especially difficult when services are addressed by domain name rather than stable IP. Standard NetworkPolicy does not select destinations by DNS name. If an application must reach SaaS endpoints with changing addresses, consider controlled egress proxies or CNI-specific extensions, and document that those mechanisms are outside portable Kubernetes NetworkPolicy semantics.

Policies should be reviewed with service ownership changes. If a backend moves to a different namespace, the old allow rule may remain while a new broad exception is added under time pressure. Periodic graph reviews can remove stale paths and keep the declared communication model understandable.

Be cautious when troubleshooting by deleting policy objects. Removing one policy may immediately expose the selected Pods to traffic from far more sources. Prefer creating a narrowly scoped temporary permit, capturing evidence, and removing the exception after the incident. Emergency changes should be auditable and time-bounded.

Policy observability is strongest when flow records include Kubernetes identity such as namespace, Pod, labels, and policy verdict. Raw IP addresses change frequently as Pods are recreated. Metadata-aware telemetry lets responders answer which workload attempted the connection and which policy allowed or denied it.

NetworkPolicy can also support multi-tenant clusters, but it is only one boundary. Tenants need RBAC, resource quotas, Pod security, secret separation, and controls over namespace creation and label mutation. Network isolation should be evaluated as part of that wider tenancy model.

When using a service mesh, understand which component originates the packet that policy sees. Sidecars, gateways, and transparent proxies can alter source or destination behavior depending on CNI integration. Test policy with the mesh enabled rather than assuming a non-mesh laboratory result will carry over.

A useful review artifact is a table of required flows: source identity, destination identity, protocol, port, business purpose, and owner. That table makes policy intent reviewable by application teams and gives security engineers a concrete basis for detecting unexpected communication.

Dual-stack clusters deserve explicit tests. IPv4 and IPv6 paths can behave differently depending on CNI support, Service configuration, and external routing. A policy review that validates only one address family may miss an unintended path on the other.

Be aware of node-local and host-network workloads. Components running with host networking do not always behave like ordinary Pods from the policy engine’s perspective. Document how your CNI treats these flows, especially for ingress controllers, DNS caches, and security agents.

When a policy change causes an outage, resist the temptation to replace it with a universal allow. Compare the failing flow with the intended communication matrix, add the smallest rule that represents the legitimate dependency, and retain the incident as evidence that the dependency was previously undocumented.

Policy reviews should include ports exposed by new application versions. A service that adds a metrics or management listener can create a path that the old policy did not anticipate. Decide whether the new listener should be reachable at all and by which identity before changing the allow rules.

Finally, use policy names and labels that describe intent. Names such as allow-frontend-to-orders are easier to audit than generic names such as network-policy-7. Readable policy reduces mistakes when several additive rules select the same Pods.

Include policy in disaster-recovery and migration testing. A restored cluster that recreates workloads but omits namespace labels or NetworkPolicy objects can come back far more open than the original. Validate isolation alongside application health so recovery does not silently remove a major security boundary.

Filed under Networking