INSIGHTS
Cloud Computing

AWS SAP-C02: Direct Connect and VPN Resilience Patterns

In this article
  1. Define the failure model before choosing the topology
  2. Use multiple Direct Connect connections for critical paths
  3. Use Site-to-Site VPN as primary, secondary, or backup connectivity deliberately
  4. Choose the AWS gateway according to scale and routing needs
  5. Engineer BGP policy for normal and failure traffic
  6. Plan encryption separately from private transport
  7. Design for asymmetric routing and stateful devices
  8. Test failover with application traffic, not just interface status
  9. Align connectivity resilience with application resilience

Hybrid connectivity is often designed as if “a Direct Connect” or “a VPN” were complete architectures. They are only components. Resilience depends on physical paths, devices, locations, virtual interfaces, gateways, tunnels, routing, failure detection, customer-premises design, and the applications that rely on the connection. A resilient pattern therefore begins with the failures the business must survive and the recovery time the network is expected to deliver.

These tradeoffs are central to SAP-C02 and even more directly related to ANS-C01. AWS Direct Connect provides private connectivity; Site-to-Site VPN provides encrypted IPsec connectivity, commonly over the internet. Strong designs often combine them, but the correct combination depends on reliability, performance, encryption, deployment time, and cost.

Define the failure model before choosing the topology

List the failures the design must tolerate: a router interface, customer edge device, Direct Connect connection, AWS device, facility, provider circuit, Direct Connect location, internet path, BGP session, tunnel, Transit Gateway attachment, or entire customer site. A design that protects against one circuit failure may still fail when both circuits share the same building entrance or provider equipment.

Translate those failures into business impact. A development office may accept hours of degraded connectivity, while a payment platform may require automatic failover with little interruption. AWS hybrid-connectivity guidance emphasizes that higher reliability increases cost, so resilience should be proportionate to workload criticality.

Record recovery targets for network failure separately from application recovery. A link may reconverge in seconds while stateful sessions take minutes to recover, DNS caches point to unavailable endpoints, or batch jobs fail and require restart. The user-visible recovery time is the metric that matters to the business.

A useful network-failure perspective reinforces the principle: redundancy is meaningful only when components do not share the same hidden failure domain.

Use multiple Direct Connect connections for critical paths

A single Direct Connect connection is a single dedicated connection between customer or partner equipment and AWS. AWS recommends additional connections when redundancy is required. The Direct Connect Resiliency Toolkit describes models including maximum resiliency, where separate connections terminate on separate devices and more than one Direct Connect location can be used to tolerate device, connectivity, and location failures.

Physical diversity must be verified with the provider. Ordering two logical circuits does not guarantee diverse fiber routes. Ask how connections enter facilities, which carriers are involved, and whether customer-premises equipment is also redundant. The end-to-end path is only as resilient as its least-diverse segment.

Use BFD where supported to improve failure detection on Direct Connect BGP sessions. Faster detection helps traffic move to surviving paths, but aggressive timers must be tested with the actual devices and provider network.

Capacity also needs diversity. Two redundant connections that are each sized for only half of peak traffic can create an outage during failover even when routing works perfectly. Decide whether each surviving path must carry 100 percent of critical load or whether the business accepts degraded capacity.

Use Site-to-Site VPN as primary, secondary, or backup connectivity deliberately

AWS Site-to-Site VPN creates encrypted IPsec tunnels between customer gateways and AWS virtual private gateways or Transit Gateway. Each VPN connection includes redundant tunnels, but customers must configure their devices to use both if they want tunnel-level resilience. VPN can serve as the main connection for smaller or rapidly deployed sites, or as backup for Direct Connect.

A VPN backup to Direct Connect is cost-effective, but AWS notes that internet-based performance is less deterministic and does not provide the same reliability characteristics as private Direct Connect. The design should therefore state whether VPN is intended to maintain full application capacity, preserve only critical traffic, or provide administrative access during a circuit failure.

Accelerated Site-to-Site VPN can be considered when internet path consistency is important and the architecture fits the service. It does not convert an internet path into Direct Connect, but it can change performance characteristics. Benchmark from the actual customer locations before depending on it for a recovery objective.

General VPN headend concepts are relevant because tunnel termination, routing, and device capacity determine how useful the backup path is under real load.

Choose the AWS gateway according to scale and routing needs

A virtual private gateway connects hybrid traffic to an individual VPC. AWS Transit Gateway provides a regional hub that can connect many VPCs, VPNs, and Direct Connect connectivity through a Direct Connect gateway. Transit Gateway is often the better fit for large multi-account environments because it centralizes routing and can support ECMP across eligible paths.

Centralization should preserve segmentation. Use separate Transit Gateway route tables or other controls so production, non-production, partners, and inspection networks do not automatically gain full connectivity. A hub that makes routing easier should not flatten the organization’s trust boundaries.

For smaller architectures, the simpler gateway may be easier to operate. The best topology is the least complex one that still satisfies scale, segmentation, and resilience requirements.

Direct Connect gateways add another scope decision by associating private or transit virtual interfaces with VPC or Transit Gateway connectivity across Regions and accounts according to supported models. Keep associations deliberate and document which network team owns route propagation and allowed prefixes.

Engineer BGP policy for normal and failure traffic

Dynamic routing is central to hybrid failover. BGP advertisements determine which prefixes are reachable and which paths are preferred. Understand AWS path-selection behavior, local customer routing policy, AS paths, communities where applicable, and how routes change when a connection or tunnel fails.

A BGP foundation is essential because a backup link is useless if routing continues to prefer the failed or black-holed path. Test route withdrawal and convergence rather than assuming the control plane will behave as the diagram suggests.

Watch for route-count differences between connectivity types. AWS hybrid-connectivity guidance notes that VPN and Direct Connect paths terminating on Transit Gateway can have different route-advertisement characteristics, which may create unexpected preferences. Keep advertisements intentional and summarize where appropriate.

Control failback as carefully as failover. When the preferred Direct Connect path returns, BGP may immediately move traffic back while applications are still stabilizing or the provider is still testing the circuit. Operational teams should know whether automatic restoration is acceptable or whether a maintenance hold is required.

Plan encryption separately from private transport

Direct Connect is private connectivity, but private does not automatically mean encrypted at the network layer. If policy requires encryption in transit, options can include Site-to-Site VPN over Direct Connect or MACsec on supported dedicated connection speeds and configurations. Application-layer TLS remains another layer but may not satisfy every network-security policy.

VPN over the public internet provides IPsec encryption by design. When both DX and VPN exist, document whether traffic is encrypted on both paths and whether the security posture changes during failover. A disaster mode should not quietly weaken a regulatory control.

Key management, tunnel parameters, certificate or pre-shared-key lifecycle, and customer-device hardening are operational responsibilities. Connectivity security is a continuous process, not only an architecture choice.

Document MTU and fragmentation expectations as well. Encapsulation from IPsec, GRE, SD-WAN, or other overlays can reduce effective payload size. A path that passes simple pings may still break large application packets or perform badly if path MTU discovery is blocked.

Design for asymmetric routing and stateful devices

Multiple paths can produce asymmetric traffic, where packets leave through one connection and return through another. Pure IP routing can tolerate this, but stateful firewalls, NAT devices, intrusion-prevention systems, and some customer appliances may not. Failure conditions can therefore break sessions even though routes technically remain available.

Map stateful components on both sides of the connection. If centralized inspection is required, make sure active and backup paths traverse compatible inspection states or that the application can tolerate session re-establishment. ECMP can improve scale, but only when the end-to-end design supports it.

Monitoring should detect route asymmetry, tunnel health, BGP state, latency, packet loss, and interface errors. The AWS network-optimization view is useful because resilience depends on observing path behavior, not merely knowing that a connection object exists.

Correlate AWS metrics with customer-router telemetry and provider status. A tunnel can look healthy from one side while drops, optical errors, or congestion exist elsewhere. End-to-end telemetry shortens the argument over which domain owns the problem and helps operations choose the correct failover action.

Test failover with application traffic, not just interface status

A successful BGP failover does not prove that the application works. Test DNS resolution, authentication, API dependencies, MTU, throughput, long-lived sessions, firewall state, and application timeouts during each planned failure. Measure how long traffic is disrupted and whether users experience errors.

Run controlled game days that disable a tunnel, circuit, router, or simulated site path. Validate both failover and failback. Some networks recover to a suboptimal route or create oscillation when the preferred path returns. Operational procedures should define how to verify stability after restoration.

Include provider escalation and out-of-band management. If the hybrid path is down, the team still needs a way to reach customer routers and communicate with carriers. The best technical redundancy can be undermined by an incident process that depends on the failed network.

Capture baseline performance before failure tests. Latency, jitter, throughput, and packet loss on normal and backup paths provide evidence for capacity decisions and make it easier to distinguish expected degraded performance from an actual fault during the exercise.

Align connectivity resilience with application resilience

Highly redundant hybrid networking cannot rescue an application that has one on-premises dependency with no failover. Conversely, an application designed to run independently in AWS may not need the most expensive hybrid topology. Architecture teams should align network availability with workload recovery objectives and data flows.

The SAA-C03 foundation of Multi-AZ design, scaling, and recovery helps place hybrid networking in context. Direct Connect and VPN should support the application’s reliability pattern, not become a separate prestige project.

Classify traffic by criticality. During degraded operation, routing or QoS policies may prioritize authentication, transaction, replication, or management traffic over bulk transfers. A backup VPN that cannot carry normal peak volume can still be valuable if the business has planned which functions remain essential.

The final topology should be documented in terms of normal path, each failure path, expected BGP state, encryption, monitoring, ownership, and escalation. Avoid designs with many redundant components but no operator who can explain which route should be active. Complexity itself can become a reliability risk.

Common patterns range from dual VPN tunnels for modest requirements, through Direct Connect with VPN backup, to multiple Direct Connect connections across locations plus independent VPN or additional DX paths for critical workloads. There is no universally correct level. AWS explicitly frames resilience as a tradeoff between failure tolerance and cost.

Resilient hybrid connectivity is successful when failures are predictable rather than surprising. Identify the failure domains, diversify them intentionally, engineer routing, secure both primary and backup paths, test with real application traffic, and keep the operational playbook current. That discipline matters more than the number of lines drawn between the data center and AWS.

Revisit the design when traffic grows, sites are added, providers change, or new AWS Regions become part of the application. A topology that was resilient at 500 Mbps may not be resilient at 5 Gbps if the backup path did not scale with it. Resilience is a maintained property, not a one-time topology decision.

Capacity calculations should use failure-state traffic, not only normal utilization. If two Direct Connect links normally split a 4 Gbps workload, each remaining path may need to carry the full 4 Gbps after a failure. The same reasoning applies to VPN backup: estimate which traffic must continue, the throughput the tunnels and customer devices can sustain, and which flows can be deferred during degraded operation.

Physical diversity should be verified with providers rather than inferred from different circuit identifiers. Two connections can terminate in different routers yet share a carrier, conduit, building entrance, or metro path. Critical designs document which failure domains are genuinely independent across customer equipment, providers, Direct Connect locations, and AWS Regions.

Transit Gateway can simplify many-site and many-VPC connectivity, but centralized routing increases the importance of route-table design. Separate propagation and association according to trust and environment, understand how VPN and Direct Connect Gateway attachments advertise routes, and avoid broad propagation that creates unintended reachability during failover.

Operational alarms should distinguish a hard outage from degradation. Rising packet loss, BGP flaps, optical errors, or sustained latency may justify moving traffic before a circuit fails completely. Define thresholds and authority for manual or automated path changes, because repeated oscillation between marginal links can be worse than staying on one degraded but stable path.

Document dependency on public internet services during a Direct Connect failure. DNS, identity providers, certificate validation, monitoring endpoints, and management tools may take different paths than application data. A backup VPN test should confirm those auxiliary dependencies as well, otherwise the primary packets may reroute successfully while operators lose the tools required to manage the incident.

Finally, keep diagrams synchronized with routing policy and circuit inventory. Incident responders should be able to see provider, location, VLAN, virtual interface, gateway attachment, tunnel, BGP peer, advertised prefixes, and expected failover preference without reconstructing the network from consoles. Accurate operational documentation reduces recovery time when several connectivity events occur together.

Filed under Cloud Computing