Elastic Load Balancing sits at the boundary between clients and application capacity, but it is more than a device that spreads requests evenly. The load-balancer type determines which protocols are understood, what routing decisions are available, how targets are health-checked, and where public, private, or inspection boundaries can be placed in the network.
For SAA-C03, the key skill is to map the application tier to the right balancing model. Application Load Balancers, Network Load Balancers, and Gateway Load Balancers solve different problems. A strong architecture uses listeners, rules, target groups, health checks, and multi-AZ placement to make service boundaries explicit rather than hiding unrelated workloads behind one endpoint.
Place load balancing at real service boundaries
Elastic Load Balancing distributes traffic to healthy targets and is a core part of resilient AWS application design. For SAA-C03, the important decision is not merely to “add a load balancer,” but to decide where one tier should be isolated from another and which protocol features each boundary requires.
An internet-facing load balancer can separate public clients from a private application fleet. An internal load balancer can decouple a web tier from an internal API tier. A Gateway Load Balancer can insert a fleet of virtual security appliances into a network path. These are different architectural roles even though they share the load-balancing family name.
The basic mechanics described in cloud load balancing are the foundation: listeners receive connections, rules select target groups, health checks remove unhealthy targets, and capacity can span Availability Zones.
An application can have more than one load-balancing layer without being overengineered if each layer has a distinct responsibility. CloudFront may distribute edge requests globally, an internet-facing ALB can route HTTP paths to services, and an internal NLB can expose a private protocol to another system. The danger is adding layers whose purpose nobody can articulate; each extra hop adds logs, timeouts, health semantics, and failure modes.
Choose ALB for HTTP-aware application routing
Application Load Balancers operate at layer 7 and understand HTTP and HTTPS. Listener rules can route based on host names, paths, headers, query strings, source IP, and other request attributes. This makes an ALB a natural front door for web applications, APIs, and microservices that need multiple target groups behind one endpoint.
Use host and path routing to represent ownership boundaries. `api.example.com` can route to one service while `app.example.com` routes to another, or `/orders/*` can reach an order service while `/catalog/*` reaches a catalog service. Keep rules understandable and avoid building an accidental application router with hundreds of fragile conditions.
TLS termination at the ALB centralizes certificate management, but backend encryption can still be required for sensitive paths. The network security model should state where plaintext is acceptable and which internal hops must remain encrypted.
ALB listener-rule priorities should be reviewed like application routing code. Overlapping host and path rules can send traffic to an unintended service, especially as microservices multiply. Use infrastructure-as-code tests that evaluate representative URLs and expected target groups. This catches route shadowing before a production deployment turns one service release into another service’s outage.
Choose NLB for transport-level requirements
Network Load Balancers operate at layer 4 for TCP, TLS, UDP, and related transport traffic. They are designed for very high performance, preserve source IP behavior in common target modes, and support use cases that need static IP characteristics or AWS PrivateLink integrations.
Do not select NLB because it sounds faster for every application. If the architecture needs path-based routing, HTTP redirects, web application firewall integration at the ALB layer, or request-aware decisions, an ALB is normally a better semantic fit. The right protocol layer matters more than a generic performance claim.
NLB can also target an ALB in supported designs, combining transport-level features such as static IPs or PrivateLink with layer-7 routing behind it. That should be used because both layers provide a needed capability, not as a default stack.
NLB static IP and Elastic IP features are useful when external partners allowlist source or destination addresses, but they should not become an excuse to expose backends directly. Keep target resources private and let the load balancer own the stable entry point. When TLS is passed through or terminated at the NLB, document where certificates and client identity are validated.
Use Gateway Load Balancer for appliance fleets
Gateway Load Balancer is designed to deploy and scale virtual appliances such as firewalls, intrusion detection systems, and deep packet inspection tools. It operates at the network layer and uses GENEVE to forward traffic to appliance targets.
This is different from putting a web application behind an ALB. The target is an inspection or networking appliance whose purpose is to process flows, often as part of centralized security architecture. Routing and symmetry become critical because stateful appliances may require both directions of a flow to traverse the same logical inspection path.
These designs connect to the broader security architecture covered by AWS security specialty topics. Network inspection should complement, not replace, workload security groups, IAM, application authentication, and service-native controls.
Gateway Load Balancer designs should account for appliance scaling independently from application scaling. A web traffic surge can overload inspection targets even if application instances are healthy. Monitor appliance flows, target health, and drop behavior, and decide whether cross-zone inspection is acceptable from both latency and data-transfer perspectives.
Build target groups around deployment units
A target group defines the backend protocol, port, target type, and health checks used by a listener rule. Create target groups around deployable service units rather than around arbitrary infrastructure collections. That makes blue/green deployment, canary traffic shifting, and service ownership easier to reason about.
Targets can be instances, IP addresses, Lambda functions for ALB, or other supported types depending on the load balancer. Containers often register task or pod IPs, while older applications may register EC2 instances. The target type should match how the platform creates and replaces compute.
Keep one target from serving unrelated applications simply to reduce the number of groups. Isolation helps health checks reflect the actual service and allows scaling one tier without coupling it to another.
Blue/green releases can use separate target groups with weighted forwarding or deployment services to move traffic gradually. Health checks should prove the green environment is ready before weight increases. Keep rollback fast by retaining the previous target group long enough to observe real traffic rather than deleting it immediately after deployment.
Design health checks for serving ability
Health checks determine whether the load balancer considers a target capable of receiving new traffic. A check should verify enough of the application to catch local failure without depending on every downstream system. If the health endpoint calls the database, payment provider, and message broker, one shared dependency outage can make all targets fail simultaneously.
Use readiness checks for traffic eligibility and separate deeper synthetic tests for end-to-end service health. A process can be alive while still warming caches or applying migrations, so a simple TCP connection may mark it healthy too early. Conversely, a target should not be removed because an optional downstream feature is degraded.
Review thresholds and intervals with recovery time in mind. Aggressive checks detect failure quickly but can also flap during transient latency. Grace periods are important during deployments so a new target has time to initialize before being judged unhealthy.
Connection draining matters for long-lived HTTP requests, streaming, and WebSocket traffic. Deregistration delay gives established connections time to finish, but a deployment that terminates targets faster than clients reconnect can still produce visible errors. Align application shutdown hooks, container stop timeouts, and load-balancer draining behavior.
Spread traffic across Availability Zones intentionally
Load balancers can use subnets in multiple Availability Zones, but resilience depends on targets being available across those zones. If every registered target lives in one AZ, the load balancer cannot manufacture multi-AZ application capacity. Deploy and scale target groups with zone balance as an explicit goal.
Cross-zone load balancing behavior differs by load balancer type and configuration. Understand whether each load-balancer node sends traffic only to local-zone targets or can use healthy targets in all enabled zones. That affects capacity planning and cross-zone data transfer.
Multi-AZ thinking connects to availability targets: load balancing is the routing mechanism, while spare capacity and independent failure domains are what make the service survive a zonal event.
Zone-level capacity should be modeled under both cross-zone enabled and disabled states. If each load-balancer node prefers local targets, a zone needs enough local capacity for its incoming traffic. If cross-zone routing is enabled, the remaining zones must absorb redirected traffic during failure. The configuration changes not only availability but also the shape of backend load.
Protect internal tiers with private load balancers
Not every service needs a public endpoint. Internal ALBs and NLBs can expose a stable address only inside the VPC or connected networks. This is useful for service-to-service traffic, shared APIs, and legacy applications that need load balancing without internet reachability.
Security groups on supported load balancers and targets should express the intended flow. An application tier can allow inbound traffic from the load balancer’s security group rather than from an entire CIDR block. This narrows the path and makes changes easier to review.
Network segmentation concepts such as multi-subnet design help separate public ingress, private application tiers, and data tiers. Load balancers should reinforce those boundaries rather than create shortcuts around them.
Private load balancers can become shared infrastructure, so ownership matters. If ten applications depend on one internal NLB, a listener change can have broad blast radius. Prefer clear service boundaries and separate change control where teams release independently. Shared endpoints are useful when they represent a real shared service, not merely an effort to reduce the number of load balancers.
Operate load balancing as part of application delivery
Monitor target response time, healthy host count, rejected connections, HTTP error codes, and backend-specific metrics together. An ALB can show elevated 5xx responses because targets are failing, while targets may look healthy if the issue occurs only on a specific route. Logs and application traces are needed to bridge the layers.
Deployment systems should register new targets, wait for health, shift traffic, and drain old targets gracefully. Deregistration delay gives in-flight requests time to complete, but very long-lived connections may need application-specific handling. Test rollbacks through the load balancer rather than assuming registration changes are instantaneous.
The advanced SAP-C02 view is that load balancing participates in multi-Region, hybrid, and multi-account architecture. At any scale, the principle stays simple: choose the correct load-balancer type, align target groups with services, make health meaningful, and preserve independent capacity across failure domains.
Capacity and quota reviews should include target count, listeners, rules, certificates, and connection patterns before large launches. Elastic Load Balancing scales automatically for most workloads, but surrounding limits and backend capacity can still become constraints. A load test that reaches only the load balancer and not realistic targets proves very little about the full application path.
(8, ‘Load balancer timeouts should also align with application behavior. An HTTP service with long-running requests can be cut off by an idle timeout even when the backend is still working, while excessively long timeouts can hold resources for abandoned clients. Streaming and WebSocket workloads deserve explicit timeout and connection-lifetime testing rather than inheriting defaults intended for ordinary web requests.’)
(8, ‘Certificate and DNS operations belong in the availability plan. A perfectly healthy ALB is useless if the certificate expired or the DNS alias was changed incorrectly. Use managed certificate renewal where possible, monitor expiration, protect DNS changes, and test failover names. Entry-point resilience includes the controls that let clients discover and trust the load balancer.’)
(8, ‘When a workload uses AWS WAF with an ALB or CloudFront, keep responsibilities clear: WAF filters application-layer requests, the load balancer selects healthy targets, and security groups constrain network reachability. Layered controls are valuable when each has a defined job. Duplicating the same broad allow/deny logic in every layer makes incident troubleshooting harder.’)
(8, ‘Review cost at the same time as availability. Cross-zone routing, idle capacity, multiple load balancers, and data processing all have charges. Consolidating unrelated services behind one giant ALB can reduce line items but increase blast radius and rule complexity. Paying for separate load balancers can be justified when it preserves ownership, security boundaries, or independent scaling.’)
HTTP/2, gRPC, TLS policies, client IP preservation, and WebSockets can influence load-balancer choice as much as raw traffic volume. Review protocol requirements before implementation so the team does not discover late that a needed client behavior is incompatible with the chosen listener or target type. Protocol fit should be decided during architecture design, not during incident troubleshooting.
Access logs should be retained long enough to support performance and security investigations. They can show request paths, status codes, target timing, and client characteristics that are not visible in aggregate metrics. Pair logs with tracing and backend request IDs so a slow or failed request can be followed from the load balancer into the service.
Treat load-balancer configuration as production code. Listener rules, target-group attributes, TLS policies, health checks, and security groups should be versioned, reviewed, and tested before promotion. Console-only changes are difficult to reproduce and can create silent drift between environments. A repeatable configuration also makes disaster recovery easier because the entry path can be rebuilt from the same source as the application.