INSIGHTS
Cloud Computing

AWS ANS-C01: BGP Design for Direct Connect

In this article
  1. Choose the virtual interface that matches the destination
  2. Design the prefix plan before tuning attributes
  3. Understand AWS route-selection order
  4. Use local-preference communities for return-path control
  5. Use AS-path prepending only after simpler controls
  6. Design public-VIF communities around propagation scope
  7. Build active/active paths only when the network can use them
  8. Combine Direct Connect with VPN deliberately
  9. Verify routing with evidence, not assumptions

AWS Direct Connect gives an organization a dedicated path into AWS, but Border Gateway Protocol determines whether that path behaves the way the architecture expects. The physical circuit can be healthy while traffic follows the wrong connection, fails over too slowly, advertises too much, or creates asymmetric return paths. Reliable Direct Connect therefore requires engineers to treat BGP policy, prefixes, communities, and failure domains as part of the application architecture rather than as provider-side plumbing.

The current AWS Certified Advanced Networking – Specialty (ANS-C01) exam explicitly covers routing between on-premises networks and AWS, including Direct Connect, routing protocols, hybrid connectivity, and network operations. Candidates should understand how public, private, and transit virtual interfaces differ, how AWS evaluates routes, and how to build active/active or active/passive designs without depending on accidental BGP behavior.

Choose the virtual interface that matches the destination

A Direct Connect connection becomes useful only after a virtual interface is created. A private virtual interface reaches VPC resources through private addressing, a public virtual interface reaches AWS public service endpoints using public addressing, and a transit virtual interface connects to one or more Transit Gateways through a Direct Connect gateway. These are different routing domains, not interchangeable ways to attach the same circuit.

Start with the destination and the intended ownership boundary. A workload VPC that needs private connectivity may fit a private VIF, while a multi-VPC network built around Transit Gateway often fits a transit VIF. Public VIFs introduce public-prefix advertisement rules and should be used when the requirement is private transport to AWS public services rather than private VPC address space. Avoid creating extra VIF types merely because the port can support them.

Redundancy planning begins here. If the organization needs independent failure domains, use separate Direct Connect locations, devices, and connections rather than placing both BGP sessions on the same physical dependency. BGP can select between available paths, but it cannot create physical diversity that the circuit design does not provide.

Design the prefix plan before tuning attributes

Route preference begins with the routes being advertised. AWS and on-premises routers apply longest-prefix match before most policy attributes, so a more-specific prefix can override a path preference that looks correct in a BGP diagram. Before adding AS-path prepending or communities, document every prefix advertised over Direct Connect, VPN, and other hybrid paths and identify where specificity differs.

Summarization reduces route-table size and policy complexity, but excessive aggregation can hide failure boundaries. If two sites advertise the same aggregate and only one has a working path to a specific subnet, traffic can be attracted to the wrong location. Use aggregates when the site can actually deliver traffic for the whole advertised range or when more-specific exceptions are deliberately advertised elsewhere.

The fundamentals in Border Gateway Protocol still apply in AWS: the protocol exchanges reachability, not end-to-end health. A prefix can remain present while the application behind it is broken. Application failover therefore needs health checks or routing controls above BGP when reachability alone is insufficient.

Route ownership should be documented alongside the prefix plan. If the networking team changes an aggregate on premises while the cloud team advertises a new more-specific prefix from another site, both changes can be individually valid and collectively wrong. Maintain one hybrid-routing source of truth that shows which location is authoritative for each prefix and which attributes are expected on primary and backup paths.

Understand AWS route-selection order

For private and transit virtual interfaces, AWS evaluates longest prefix first. When prefix length is equal, Direct Connect local-preference communities can influence return-path preference. AS_PATH length can then distinguish otherwise comparable routes, and MED is lower in the decision sequence and is not the preferred design tool for most Direct Connect path control. Equal paths can be used for ECMP when the relevant attributes match.

This order matters because engineers sometimes prepend AS paths and expect that to override a more-specific route. It will not. A /24 advertised on the backup path can beat a /16 on the intended primary path regardless of prepending. Troubleshooting should therefore start with prefix length and the actual routes AWS receives before focusing on advanced attributes.

The existing BGP attributes model is useful, but Direct Connect adds AWS-defined community behavior. Treat those communities as part of the provider edge policy and document them alongside on-premises route maps. A route policy is maintainable when another engineer can explain why the selected path wins without relying on trial-and-error changes.

Be careful when the same prefix is advertised from Direct Connect locations associated with different AWS Regions. Region association influences default path preference behavior, and custom local-preference communities can override that intent. If a multi-Region network needs deterministic egress, test the exact combinations of location, Region association, and community rather than assuming all Direct Connect locations are equivalent.

Use local-preference communities for return-path control

Direct Connect supports local-preference communities for private and transit VIF advertisements. The documented values include low, medium, and high preference. Applying the same preference to comparable paths can support active/active behavior, while assigning a higher preference to the intended primary connection and lower preference to the standby can produce an active/passive return path from AWS.

These communities influence AWS-to-on-premises traffic. They do not automatically control the on-premises router’s path toward AWS, so inbound and outbound policy must be designed separately. It is possible for on-premises traffic to choose one circuit while AWS return traffic chooses another. Asymmetric routing is not inherently wrong for stateless routing, but stateful firewalls, NAT, or monitoring appliances may require symmetry.

Use the communities consistently across all prefixes that should share the same behavior. A mixed policy where some prefixes carry a high-preference tag and others rely on default behavior can create difficult partial failures. Build a route-policy test that verifies received communities and expected best paths before the change is treated as production-ready.

Use AS-path prepending only after simpler controls

AS-path prepending is commonly used to make a route less attractive, but it is not the first tool to reach for in Direct Connect. It only matters after prefix length and higher-priority Direct Connect preference logic have been evaluated. Excessive prepending can also make the design opaque when different devices add different numbers of copies for historical reasons.

Use prepending when you need on-premises or provider-side BGP selection that cannot be expressed more cleanly with prefix design or supported community controls. Keep the policy small and documented. A route map that prepends the backup path by a fixed amount is easier to reason about than a collection of exceptions attached to individual prefixes.

The comparison in OSPF versus BGP highlights why this is different from an interior routing protocol. BGP policy is intentionally administrative. The “shortest” path is not necessarily the physically shortest link; it is the path that wins the configured policy and advertised attributes.

Design public-VIF communities around propagation scope

Public virtual interfaces have different community semantics from private connectivity. AWS supports scope communities that influence whether advertised public prefixes are propagated locally, continent-wide, or globally within the AWS network. Public VIF design also requires ownership and registration of the public prefixes being advertised and must follow AWS validation and source-filtering rules.

Do not use a public VIF to solve a private-VPC routing requirement. Public AWS service endpoints and private VPC CIDRs are different address domains. Where private workloads need controlled access to AWS services, evaluate private connectivity options such as VPC endpoints alongside routing requirements instead of forcing traffic through a public advertisement model.

Public-prefix scope is also an availability decision. A local propagation scope can reduce the number of AWS Regions that learn the route, while broader scope improves reachability but increases the effect of an incorrect advertisement. Treat BGP communities as deployment controls with peer review and rollback, not as harmless metadata.

ECMP also changes monitoring. A circuit can carry only a fraction of the flows and still be essential to aggregate capacity. Monitor per-connection utilization, errors, BGP state, and traffic distribution so one degraded link is not masked by healthy peers. An active/active design should alert when it silently becomes active/partial long before total capacity is exhausted.

Build active/active paths only when the network can use them

Equal-cost multipath can distribute traffic across multiple comparable Direct Connect paths. This can increase aggregate throughput and use redundant circuits efficiently, but only when the downstream network and stateful devices can tolerate traffic using either path. If one site has a smaller firewall cluster, different NAT behavior, or different application reachability, ECMP can expose those differences continuously rather than only during failover.

Active/passive designs are often easier to operate because the normal path is deterministic. The tradeoff is that the backup circuit may be underused and faults can remain hidden until failover. Test it regularly. A backup path that is “up” at BGP level but blocked by a forgotten firewall rule is not a backup.

The AWS Certified Solutions Architect – Professional (SAP-C02) perspective is useful here: routing choices should support application resilience objectives. Choose active/active when the full stack is designed for parallel paths, and choose active/passive when deterministic state and controlled failover are more important than using every circuit all the time.

Consider failure domains above the routing protocol. A primary and backup Direct Connect circuit that share a colocation facility, provider aggregation router, or customer edge device may fail together. True resilience usually requires diversity across physical location, carrier or logical provider path, and customer equipment as well as separate BGP sessions. The route policy can only choose among paths that remain physically available.

Combine Direct Connect with VPN deliberately

Site-to-Site VPN is frequently used as backup for Direct Connect, but the failover design must account for route specificity, BGP attributes, tunnel health, and traffic capacity. If the VPN advertises more-specific routes than Direct Connect, it can become primary unintentionally. If the backup link is much smaller than the dedicated circuit, a successful routing failover can still produce an application outage through congestion.

Test failover under realistic load and verify both directions. Shut or withdraw the primary route in a controlled window, confirm that the expected VPN route becomes active, and observe application latency and stateful devices. The general IPsec principles around encryption and tunnel behavior matter, but hybrid AWS operations add route policy and managed-gateway behavior to the troubleshooting path.

When multiple Direct Connect locations, VPN tunnels, and Transit Gateways are involved, draw the route advertisements for each failure scenario. Diagrams that show only physical links are incomplete. The operational diagram should show prefixes, BGP sessions, communities, expected best paths, and where stateful inspection occurs.

Automate route validation where possible. Periodically compare received prefixes and attributes with an expected policy and alert on unexpected more-specific routes, missing communities, or new AS paths. Hybrid environments change frequently, and a routing check performed only during the original deployment cannot protect against a later data-center migration or carrier change that alters advertisements.

Verify routing with evidence, not assumptions

Before declaring a Direct Connect design complete, capture the BGP state, received and advertised prefixes, communities, AS paths, and route-selection results on both sides. Monitor session state and traffic, but also monitor application reachability. A stable BGP session can coexist with a black hole caused by a route table, security control, or downstream device.

Change one policy dimension at a time during troubleshooting. Confirm prefix length, then preference communities, then AS path, then lower-priority attributes. Avoid simultaneous changes on both peers because a successful outcome will not reveal which change mattered. For production changes, define expected route tables before execution and compare them with the post-change state.

Document rollback in the same language as the forward change. If a new community or route map produces unexpected behavior, operators should know which exact advertisement to restore and how to confirm that the former best path returned. Clear rollback is especially important because BGP changes can affect many applications at once without changing a single workload resource.

Maintenance procedures should state whether a circuit is drained by lowering preference, withdrawing prefixes, or shutting the BGP session. Each method has different visibility and convergence effects. Graceful draining with routing policy can move traffic before hardware work begins, reducing session loss and proving that the backup path can carry production load while operators still have time to react.

Keep carrier and AWS maintenance notifications in the same operational calendar as routing changes. If a planned policy change overlaps provider work on the backup circuit, the network can accidentally lose both designed paths. Resilience depends on coordination as well as topology, especially when separate teams manage the customer edge, Direct Connect location, and AWS transit architecture.

Use maintenance windows to validate convergence metrics as well as final path selection. Record how long route withdrawal, best-path recalculation, and application recovery actually take. Those measurements turn abstract routing preferences into usable recovery objectives and help teams decide whether BGP failover alone is fast enough for the applications on the link.

BGP design for Direct Connect is ultimately a policy-engineering problem. Use the correct VIF, advertise only legitimate prefixes, choose preference controls intentionally, preserve physical diversity, and test every backup path. When route selection is explainable and verified, Direct Connect becomes a predictable hybrid transport rather than a high-speed circuit whose behavior depends on inherited BGP folklore.

Filed under Cloud Computing