INSIGHTS
Networking

AWS SAA-C03: NAT Gateway vs NAT Instance

In this article
  1. Start with the purpose of outbound translation
  2. Prefer managed NAT gateways for standard egress
  3. Understand zonal NAT gateway resilience
  4. Account for Regional NAT gateways in current designs
  5. Use NAT instances only when their flexibility is required
  6. Compare protocol and feature support
  7. Reduce NAT demand with private service access
  8. Monitor connection and port pressure
  9. Make the decision from availability, control, and cost

NAT remains a common part of AWS IPv4 architectures because private workloads often need outbound access to software repositories, partner APIs, or the public internet without accepting unsolicited inbound connections. The long-standing comparison between a managed NAT gateway and a self-managed NAT instance still matters, but current AWS networking adds a newer wrinkle: NAT gateways can now be zonal or Regional for supported public NAT designs.

This article updates the SAA-C03 decision model for modern AWS. It compares managed and instance-based NAT, explains the availability and operational differences between zonal and Regional managed gateways, and shows how route tables, VPC endpoints, monitoring, and connection behavior affect both resilience and cost.

Start with the purpose of outbound translation

NAT lets resources with private IPv4 addresses initiate connections to destinations outside their private network while hiding those private addresses from the remote side. In AWS, private subnets commonly use NAT when workloads need outbound internet access but should not receive unsolicited inbound internet connections. That model is part of the networking knowledge expected around SAA-C03.

The networking concept itself is explained in network address translation fundamentals. The AWS design decision is how translation is delivered: a managed NAT gateway, a self-managed NAT instance, or in some architectures a design that avoids internet NAT by using VPC endpoints and private service connectivity.

Modern AWS also adds an important distinction inside the managed option: zonal NAT gateways and Regional NAT gateways. That means a current architecture discussion should not assume every NAT gateway is tied to exactly one subnet and one Availability Zone.

NAT should also be separated from inbound publishing. A public NAT gateway enables outbound connections from private IPv4 resources; it is not a reverse proxy and cannot be used by internet clients to initiate connections to those private instances. Public services should use load balancers, API gateways, or other intentional ingress mechanisms rather than trying to reuse an egress component.

Prefer managed NAT gateways for standard egress

AWS recommends NAT gateways over NAT instances for most workloads because the managed service provides higher availability, greater bandwidth, and less administration. Teams do not patch an operating system, replace a failed EC2 instance, configure source/destination checks, or engineer their own failover just to provide outbound translation.

Managed NAT reduces the number of infrastructure concerns that can interrupt egress. That is especially valuable because NAT failure often affects software updates, external APIs, package repositories, identity endpoints, and other dependencies that application teams may not realize share the same path.

Operational simplicity is not free. NAT gateways have hourly and data-processing charges, and data transfer can rise when routes cross Availability Zones. Architecture should therefore combine availability and cost rather than choosing the managed service without considering topology.

Managed NAT gateways scale far beyond what many individual EC2 NAT instances can provide, but architecture should still watch destination-specific connection limits and packet behavior. High-throughput clients talking to a small set of remote endpoints can exhaust ports before aggregate bandwidth becomes the problem. Multiple gateway addresses or workload distribution may be needed.

Understand zonal NAT gateway resilience

A zonal NAT gateway is created in a specific Availability Zone and is redundant within that zone. If private subnets in several zones all route through one zonal NAT gateway, they share that zone as an egress dependency. The classic high-availability pattern is to create a NAT gateway in each active AZ and route private subnets to the NAT gateway in the same AZ.

Local-zone routing also reduces unnecessary cross-zone data transfer. Separate private route tables per AZ make the intended path explicit. If a zone fails, the surviving zones keep their own NAT path instead of depending on a gateway in the failed zone.

Do not confuse redundancy inside the NAT service with multi-AZ architecture around it. A zonal NAT gateway can be highly available within its zone while still being the wrong shared dependency for a multi-zone application.

A zonal design should also consider failure recovery of Elastic IP allowlists. If partners allowlist the NAT gateway’s EIP, replacing gateways or moving traffic across zones can change source addresses unless the architecture deliberately preserves or pre-registers them. Availability planning that ignores partner allowlists can turn a technically healthy failover into an application outage.

Account for Regional NAT gateways in current designs

Regional NAT gateways are a newer managed option that can provide public NAT across multiple Availability Zones from one Regional resource. In automatic mode, AWS manages the IP addresses and expands the NAT gateway into zones as workload network interfaces appear. This simplifies the per-AZ deployment and route-management pattern for supported public NAT use cases.

Regional NAT does not eliminate all planning. Expansion into a new Availability Zone is not instantaneous, so rapid creation of workloads in a previously unused zone can temporarily use another zone until expansion completes. Teams should also understand public IP management, pricing, monitoring, and conversion impact before replacing existing zonal gateways.

Private NAT requirements are a key exception: current AWS guidance recommends zonal availability mode for private NAT use cases. The word “Regional” should therefore be treated as a capability with defined boundaries, not as a universal replacement for every NAT gateway.

Regional NAT automatic expansion reduces manual per-AZ deployment, but teams should not assume it removes cross-zone traffic instantly when a new zone is first used. AWS documents that expansion can take time, during which traffic from the new zone may be processed through an existing zone. Capacity launches and zone expansion should therefore be staged when strict locality or transfer-cost control matters.

Use NAT instances only when their flexibility is required

A NAT instance is an EC2 instance configured to forward translated traffic. Because the customer controls the operating system and networking stack, it can support custom packet processing, specialized logging, proxies, port forwarding, or software that a managed NAT gateway does not expose.

That flexibility carries operational responsibility. The team must size the instance, patch it, monitor CPU and network performance, disable source/destination checks, configure routes, and implement failover. Bandwidth is limited by the chosen instance type, and a single instance is a single point of failure unless additional automation is built.

Use a NAT instance because a requirement needs instance-level control, not because it appears cheaper in a small spreadsheet. The engineering time and failure risk are part of the total cost.

NAT instances may still appear attractive for packet capture or transparent proxy requirements. If that flexibility is necessary, build them as a managed fleet with Auto Scaling, immutable images, health checks, and tested route failover rather than as a pet EC2 instance. The operational model should be engineered with the same seriousness as any other critical network appliance.

Compare protocol and feature support

NAT gateways support common outbound TCP, UDP, and ICMP traffic, and current public NAT designs can also participate in NAT64 with DNS64 for IPv6-to-IPv4 communication. NAT instances use the EC2 networking stack and can be customized more deeply when unusual protocols or middlebox behavior is required.

Security-group behavior also differs. NAT instances can have security groups because they are EC2 instances. NAT gateways do not use security groups in the same way, so traffic control is expressed with subnet NACLs, routing, workload security groups, and destination service policy.

Do not mix NAT with DHCP concepts. The distinction in DHCP versus NAT is useful: DHCP assigns network configuration to hosts, while NAT translates addresses across a traffic path. They solve different problems even though both appear in network setup discussions.

NACLs around NAT paths require special care because they are stateless and must allow ephemeral return traffic. Security groups on source workloads remain stateful and usually provide a clearer application policy. Troubleshooting teams should resist opening broad NACL ranges until they have traced the actual forward and return path.

Reduce NAT demand with private service access

Many workloads send large volumes of traffic to AWS services such as S3, DynamoDB, ECR, Systems Manager, or other APIs. VPC endpoints can keep supported traffic on private AWS networking and avoid sending it through an internet NAT path. This can improve security and reduce NAT processing and transfer charges.

Gateway endpoints for supported services and interface endpoints powered by PrivateLink have different routing and pricing models. Use them where the service, traffic volume, policy, and operational model justify them. An endpoint added for every possible service can become expensive and difficult to govern.

This is a broader topology optimization problem, similar to the principles in AWS network optimization. The best NAT architecture sometimes processes fewer bytes because traffic that never needed the public internet is removed from the NAT path entirely.

For S3 and DynamoDB, gateway endpoints can remove high-volume traffic from NAT without hourly endpoint charges, while interface endpoints for other services have per-hour and data-processing costs. The endpoint decision is therefore a separate cost model. Compare endpoint charges with NAT processing, security benefits, and traffic volume instead of assuming every endpoint automatically saves money.

Monitor connection and port pressure

NAT devices maintain translation state for connections. High connection rates to the same destination can exhaust available source ports or hit connection limits even when total bandwidth looks modest. Managed NAT gateways can use multiple IP addresses to increase the number of concurrent connections available to a destination.

CloudWatch metrics such as connection attempts, errors, idle timeouts, packet counts, and bytes can reveal pressure. Application connection pooling and keepalive behavior can also affect state consumption. A service that opens a new connection for every small request may create more NAT load than a service that reuses connections safely.

Troubleshooting should include route tables, subnet association, DNS resolution, NACLs, target destination behavior, and the NAT metrics. A failed outbound connection is not always a NAT capacity problem.

Large fleets should monitor ErrorPortAllocation and connection metrics because each translated connection needs source-port state. Connection pooling, HTTP keepalive, and sane client timeouts can reduce churn. A noisy client that repeatedly opens and abandons connections can create NAT pressure even when business transaction volume is moderate.

Make the decision from availability, control, and cost

Managed NAT gateways are the default for standard outbound connectivity because they remove patching and failover work. Zonal gateways fit designs that want explicit per-AZ paths or private NAT, while Regional NAT gateways can simplify public egress across zones. NAT instances remain an exception for custom networking behavior.

Subnet architecture should make the choice visible. The design ideas in multi-subnet network segmentation help separate public ingress from private application tiers, with route tables showing which private networks use NAT, endpoints, or no external egress at all.

At SAP-C02 scale, NAT is also a multi-account and hybrid-networking concern. Centralized egress can reduce the number of gateways but can increase routing complexity, blast radius, and data-processing paths. Choose the topology that meets failure-domain, governance, and cost requirements rather than assuming centralization or per-VPC NAT is always best.

Centralized egress through Transit Gateway and shared network accounts can simplify inspection and allowlisting, but it also concentrates failure and cost. Decentralized NAT keeps egress closer to each VPC but multiplies gateway count. At scale, model both topologies with data volume, inspection requirements, account boundaries, and recovery behavior before choosing.

(8, ‘For IPv6-first workloads, NAT64 and DNS64 can provide access to IPv4-only destinations, but this is a specific translation pattern rather than a reason to keep IPv4 design unchanged. Evaluate whether dual-stack endpoints, egress-only internet gateways, or native IPv6 service access can reduce dependence on IPv4 NAT over time. Address-family strategy affects both cost and future scalability.’)

(8, ‘Logging requirements can influence the managed-versus-instance decision. NAT gateways expose CloudWatch metrics and VPC Flow Logs around interfaces, but they do not provide an operating system where arbitrary packet-capture tools can be installed. If regulations or troubleshooting require deeper packet inspection, use dedicated inspection architecture rather than assuming a NAT instance is the only place to collect evidence.’)

(8, ‘Route changes during NAT migration should be treated as production changes. Updating the default route resets or redirects active flows, and converting zonal gateways to a Regional model can change source addresses or connection state. Schedule maintenance, communicate partner allowlist changes, and verify application retry behavior before altering a busy egress path.’)

(8, ‘Cost models should include hourly gateway charges, processed bytes, cross-zone transfer, endpoint alternatives, and the EC2 plus operations cost of NAT instances. A low-volume development VPC and a high-throughput production platform may reach different conclusions. Standardize the calculation method even when the resulting topology differs.’)

High availability does not mean every VPC needs the same NAT pattern. A workload with only private AWS-service traffic may need no internet NAT, a small dev VPC may accept one zonal gateway, and a large production platform may choose Regional NAT or per-AZ zonal gateways. Standardize the decision criteria so the topology matches risk instead of copying one template everywhere.

Review egress destinations as a security control. DNS filtering, network firewalls, proxies, or application allowlists may be needed when private workloads should not reach arbitrary internet hosts. NAT hides source addresses; it does not decide which destinations are trustworthy.

Finally, document which subnets use which NAT path and why. Route tables are easy to inspect individually but difficult to understand as a system when dozens of VPCs exist. A small egress inventory showing gateway type, Availability Zone behavior, source IPs, major destinations, and owner can shorten both incident response and cost reviews.

Filed under Networking