INSIGHTS
Infrastructure & Systems

AWS ANS-C01: Hybrid Connectivity Troubleshooting

In this article
  1. Define the failed flow precisely
  2. Check the physical and tunnel state before routing policy
  3. Prove the control plane route in both directions
  4. Trace Transit Gateway associations and propagations
  5. Verify VPC routes, security groups, and network ACLs
  6. Troubleshoot DNS separately from IP connectivity
  7. Use flow evidence to find the first missing hop
  8. Test failover paths before the incident
  9. Change one variable and preserve evidence

Hybrid AWS outages are difficult because the failure can sit in several administrative domains at once: an on-premises router, a Direct Connect virtual interface, a VPN tunnel, a Transit Gateway route table, a VPC route table, a security group, a network ACL, DNS forwarding, or a stateful firewall. Effective troubleshooting is therefore less about memorizing commands and more about following the packet path in a strict order until the first broken assumption is found.

The current AWS Certified Advanced Networking – Specialty (ANS-C01) exam includes maintaining hybrid routing and connectivity and analyzing network traffic to troubleshoot AWS and hybrid networks. A strong troubleshooting method separates control-plane evidence from data-plane evidence, checks both directions, and avoids changing multiple routing and security components before the failure is localized.

Define the failed flow precisely

Start with a five-tuple and a direction: source IP, destination IP, protocol, source port where relevant, destination port, and which side initiates the connection. “The data center cannot reach AWS” is too broad. One subnet may fail while another works, TCP may fail while ICMP works, or a DNS name may resolve to the wrong address while routing is healthy.

Record the expected path before investigating the actual path. Identify the on-premises gateway, Direct Connect or VPN attachment, Transit Gateway route table, VPC attachment, VPC route table, destination ENI, and return path. If a centralized firewall or NAT device is involved, place it explicitly in the diagram. The first useful troubleshooting question is not “what changed?” but “which components must this exact packet traverse?”

Scope time and blast radius as well. A failure that began immediately after a route-policy change suggests a different search order from an intermittent packet-loss issue that affects only large transfers. The general lessons in network failure analysis apply: define symptoms, isolate layers, and collect evidence before applying remediation.

Check the physical and tunnel state before routing policy

For Direct Connect, verify the physical connection and virtual interface state. A down BGP session on an otherwise healthy port means route exchange is not occurring. For Site-to-Site VPN, inspect both tunnels, IKE/IPsec state, and the customer gateway. Do not spend time debugging VPC route tables when the hybrid attachment is not actually exchanging routes.

Physical health does not prove capacity health. Error counters, optical issues, or congestion can create loss while links remain up. A VPN backup may establish correctly but be undersized for the traffic that shifted from a high-bandwidth Direct Connect circuit. The distinction between bandwidth and throughput matters during failover because the theoretical link rate does not show whether the path is currently delivering acceptable performance.

Check redundant links separately. One Direct Connect virtual interface or one VPN tunnel can be healthy while policy keeps using a degraded path. If an architecture is supposed to fail over automatically, verify which route is actually selected rather than assuming the backup took over because its status is green.

Prove the control plane route in both directions

Once connectivity to the hybrid edge is healthy, inspect route advertisements. Confirm the on-premises prefix is received by AWS and the AWS prefix is received on premises. For BGP-based designs, compare prefix length, Direct Connect local-preference communities, AS paths, and route filters. For static VPN designs, verify the exact networks configured on both sides.

Return-path mistakes are common. The forward route may send traffic over Direct Connect while AWS returns over VPN, or two Transit Gateway route tables may select different attachments. Asymmetry can work through ordinary routers but fail through stateful firewalls and NAT. Use the concepts in BGP routing to explain why the selected route wins instead of tweaking attributes until traffic happens to work.

Watch for more-specific routes and stale propagation. A /24 learned from one attachment overrides a /16 summary from another. Transit Gateway route tables can contain static routes, propagated routes, blackhole routes, and attachment associations that differ from the VPC route table. Troubleshoot each routing domain independently and then connect the results.

Direct Connect gateway associations and allowed-prefix behavior deserve their own check in Transit Gateway designs. A route can be present at the customer router but absent from the Transit Gateway path because the Direct Connect gateway is not associated as expected or the allowed prefix set does not include it. Keep that boundary visible in the troubleshooting diagram instead of treating Direct Connect and Transit Gateway as one opaque service.

Trace Transit Gateway associations and propagations

Transit Gateway is not one global route table unless the design makes it one. Each attachment is associated with a route table for lookup, and routes from attachments can be propagated into one or more route tables. A VPC attachment may therefore send traffic into a route table that does not contain the expected Direct Connect gateway or VPN route even though another Transit Gateway route table does.

For every attachment in the failed flow, document its association and the propagations that populate the relevant table. Confirm that the destination prefix resolves to the intended attachment and that no blackhole route overrides it. In segmented designs, a missing propagation may be intentional security isolation rather than a defect. Fixing it without understanding the segmentation policy can create unauthorized connectivity.

If centralized inspection is involved, check appliance mode on the inspection VPC attachment and verify symmetry through the same stateful appliance path. Route correctness on the Transit Gateway does not guarantee that the firewall sees both directions. The packet path must be consistent with the inspection architecture.

For managed services, identify whether the destination is reached through a public endpoint, gateway endpoint, interface endpoint, load balancer, or ordinary ENI. Each option changes the path. A VPC endpoint policy or private DNS setting can block a request even when the subnet route and security group are correct. Classify the destination type before assuming every AWS address behaves like an EC2 instance.

Verify VPC routes, security groups, and network ACLs

After the packet reaches the correct VPC attachment, the VPC route table must send it toward the destination subnet, endpoint, gateway, or network interface. Check the route table associated with the specific subnet, not the table you expected the subnet to use. Large environments often contain similarly named route tables left from earlier designs.

Then check security groups and network ACLs. Security groups are stateful, while network ACLs are stateless and require rules for both directions. A route can be perfectly correct while an ephemeral return port is blocked by the subnet ACL. Conversely, an overly broad security group can hide an incorrect network segmentation design during testing and later become a security problem.

Reachability tools can help reason about AWS-side configuration, but interpret their scope correctly. They can analyze supported AWS network paths and identify configuration blockers; they do not prove the external carrier, customer router, or application itself is healthy. Use them to narrow the AWS portion of the path, then correlate with on-premises evidence.

Troubleshoot DNS separately from IP connectivity

Hybrid incidents are frequently reported as network failures when the actual problem is DNS. Test the destination by IP and by name. If IP connectivity works but the name fails, inspect Route 53 Resolver inbound or outbound endpoints, resolver rules, conditional forwarders, private hosted-zone associations, and the on-premises resolver configuration.

Check which DNS server the client actually queried and which answer it received. Split-horizon names can return different addresses depending on the resolver path. Cached answers can persist after a routing or endpoint change. DNS troubleshooting should capture the query name, record type, response, TTL, and resolver used rather than relying on a browser error.

Hybrid DNS and hybrid routing must agree. A private name that resolves to an RFC1918 address is useful only if that prefix is reachable over the intended Direct Connect or VPN path. Conversely, a public answer may send traffic to the internet even though a private hybrid route exists. Name resolution is part of path selection.

Do not overlook MTU and fragmentation. Direct Connect, VPN, Transit Gateway, and intermediate appliances can support different maximum transmission behavior. Small pings may work while large application packets fail or retransmit. If the symptom depends on packet size or only appears under encrypted tunnels, test path MTU and TCP MSS rather than continuing to change routing policy.

Use flow evidence to find the first missing hop

VPC Flow Logs can show whether packets reached an ENI or subnet boundary and whether network controls accepted or rejected them. Firewall logs, load balancer logs, and application logs can move the investigation further up the stack. On premises, interface counters, firewall session tables, packet captures, and routing tables provide the corresponding evidence.

Correlate timestamps and source addresses carefully. NAT can change the source before the packet reaches AWS, so the address seen in Flow Logs may not be the address the application team expected. If a centralized inspection VPC performs NAT, capture the before-and-after addresses in the path diagram. A packet capture on both sides of a suspected boundary is often faster than reading hundreds of configuration lines.

The operational goal is to identify the first place where expected evidence disappears. If the customer router sends the packet but the Direct Connect edge never learns the route, stay at the hybrid boundary. If AWS Flow Logs show accepted traffic on the destination ENI but the application never responds, stop changing route tables and investigate the host or service.

Application dependencies can expose a partial hybrid failure. One service might reach its database on premises but fail to contact DNS, identity, or a licensing server over a different prefix. Build synthetic tests for critical dependency paths rather than using a single “ping the data center” probe. The hybrid network is healthy only when the required application paths are healthy.

Test failover paths before the incident

Redundant hybrid connectivity should be exercised during controlled maintenance. Withdraw the preferred BGP route, disable a VPN tunnel, or isolate a test prefix and verify that traffic uses the documented backup. Confirm application capacity and return-path symmetry, not just that a BGP session remains established. A failover design exists only after it has been observed working.

Be careful with route timers and stateful sessions. A path may converge quickly while existing TCP sessions reset because the return path changes. Applications with long-lived connections can experience a larger disruption than routing convergence metrics imply. Measure both network restoration and service restoration.

The AWS Certified Solutions Architect – Professional (SAP-C02) view is useful because hybrid connectivity is part of resilience architecture. The backup path must satisfy throughput, security, and operational requirements, not merely exist as a line on the diagram.

Record the incident path after recovery. Update diagrams, route inventories, and runbooks with the component that actually failed and the evidence that proved it. Hybrid outages often recur when the fix is applied but the mental model stays wrong. A short post-incident network map can be more valuable than a long timeline if it corrects how teams understand the path.

Change one variable and preserve evidence

Hybrid incidents become harder when teams simultaneously add routes, loosen security groups, restart VPNs, and alter BGP policy. If the flow begins working, nobody knows which change solved it or whether a new exposure was introduced. Make one scoped change, capture the before-and-after evidence, and revert experiments that are not part of the final design.

Maintain runbooks that list the exact consoles, commands, logs, route tables, and expected states for critical hybrid paths. The broad VPN headend model can help teams understand termination, but AWS hybrid troubleshooting also requires cloud route-domain knowledge. Runbooks should bridge both worlds so the incident does not stall between network and cloud teams.

Where possible, automate evidence collection before remediation. Snapshot route tables, attachment states, BGP summaries, Flow Log queries, and relevant health metrics into the incident record. That preserves a before-state even if an emergency change restores service and removes the original symptom.

Quotas and route scale can create less obvious failures. A new route may not appear because a service quota or route-table limit has been reached, even though the configuration request succeeded elsewhere in the workflow. Check relevant quotas when the environment has grown rapidly or when only recently added prefixes are missing while older connectivity remains stable.

Finally, distinguish reachability from authorization at the application edge. A TCP connection that reaches a database listener but is rejected by identity policy is not a network outage. Likewise, an HTTPS 403 response proves more of the path is working than a timeout. Preserve the exact error returned by the application because it tells the network team how far the request progressed.

When packet captures are available, compare sequence numbers and retransmissions rather than only checking whether a SYN appeared. Repeated retransmissions, selective loss, or one-sided FIN/RST behavior can reveal MTU, firewall, or application issues that a simple reachability test misses. Preserve captures from both sides of the suspected boundary when possible.

A disciplined sequence makes hybrid troubleshooting predictable: define the flow, verify the physical or tunnel edge, prove the control-plane routes, trace Transit Gateway, inspect VPC routing and security, validate DNS, then use packet and flow evidence to isolate the application boundary. That method reduces guesswork and prevents a local symptom from triggering unsafe changes across the entire hybrid network.

Filed under Infrastructure & Systems