OSPF troubleshooting becomes much faster when the engineer distinguishes adjacency formation from link-state database behavior. A router can fail to become a neighbor, become a neighbor but never reach Full, reach Full but lack an expected LSA, or have the correct LSA while a different route wins in the routing table. Those are different failure classes and require different evidence.
The current 300-410 ENARSI v1.1 blueprint covers OSPFv2 and OSPFv3 neighbor relationships, authentication, network types, area types, router roles, virtual links, and path preference. 350-401 ENCOR v1.2 provides the core multi-area and adjacency foundation. A disciplined workflow follows the protocol state from Hello packets to LSDB synchronization, route calculation, RIB installation, and forwarding.
Use the neighbor state to narrow the failure domain
The OSPF state machine provides diagnostic information. Down means no valid Hello relationship exists. Init means a Hello was received but the local router did not see itself listed in the neighbor’s Hello. Two-Way confirms bidirectional Hello communication. ExStart and Exchange involve database descriptor negotiation and LSDB summary exchange. Loading requests missing LSAs, and Full means the databases have synchronized for that adjacency.
Do not treat every non-Full state as identical. The correct next check depends on where progress stopped. A Down or Init problem points toward interface participation, multicast, timers, ACLs, or one-way communication. ExStart or Exchange points toward DBD exchange and commonly MTU. Loading points toward requested LSA synchronization.
Capture the state on both routers. One side can report Exchange while the other reports ExStart, and the asymmetry itself is evidence. Repeated transitions also matter because a neighbor that flaps through Full is different from one that never leaves Init.
Check Hello parameters before changing routing policy
OSPF neighbors need compatible area IDs, Hello and Dead intervals, authentication, network type expectations, and addressing on the shared link. A stub-area mismatch can also prevent adjacency. Router IDs must be unique. Verify the actual interface-level values rather than assuming process-wide configuration inherited as intended.
If no neighbor appears, confirm OSPF is enabled on the interface and the interface is not passive. Then verify that protocol 89 traffic is not blocked and that multicast Hellos reach the peer where the network type uses them. An IP ping can succeed while OSPF packets are filtered.
Authentication failures can be subtle when keys or lifetimes differ. Check the configured method, key IDs, secrets, and time where time-based keys are used. Avoid disabling authentication merely to make the adjacency come up unless the test is controlled and the secure configuration is restored immediately.
Check bidirectional reachability with the same rigor as adjacency parameters. OSPF Hellos can be multicast while later database exchange may use unicast depending on network type. An ACL, asymmetric path, or layer-2 defect can therefore allow the neighbors to discover each other but block progress later in the state machine. Packet capture on both sides can show exactly which OSPF packet is the last successful exchange.
Recognize when Two-Way is normal on broadcast networks
On a broadcast segment, routers do not normally form Full adjacency with every other router. They form full adjacencies with the Designated Router and Backup Designated Router, while other DROTHER neighbors can remain in Two-Way. Seeing Two-Way is therefore not automatically a fault.
Check the neighbor role and interface network type before trying to force the state to Full. If no DR or BDR can be elected because every router has priority zero, or if the network type is mismatched, the expected adjacency pattern can differ. Understand the topology first.
Unnecessary attempts to “fix” a healthy Two-Way state can destabilize the segment by changing priorities or network types. Use the OSPF design expectation as the reference: which routers are supposed to become adjacent on this network, and which are simply neighbors?
Suspect MTU and DBD exchange when neighbors stick in ExStart or Exchange
ExStart and Exchange are the stages where routers establish a primary/subordinate relationship and exchange database descriptor packets. An MTU mismatch is a classic reason the process stops there. The router advertising the larger interface MTU can send DBD information the neighbor refuses, causing repeated negotiation without database synchronization.
Verify interface MTU on both sides and test the path with appropriately sized packets and the Don’t Fragment bit. Cisco provides an OSPF MTU-ignore option, but using it should be an exception for a design that legitimately requires different interface MTUs. In ordinary routed Ethernet links, correcting the mismatch is safer than suppressing the check.
Other causes can include duplicate router IDs, broken unicast communication during DBD exchange, NAT effects, or unexpected sequence behavior. The state narrows the search, but it does not justify assuming every ExStart problem is MTU without evidence.
Investigate Loading as an LSA request and database problem
In Loading, the router has identified LSAs it needs and sends Link State Requests to its neighbor. If it never reaches Full, inspect which LSA is being requested and whether the neighbor can supply a valid copy. Corrupted, inconsistent, or repeatedly rejected LSAs can keep the adjacency from completing.
Check logs for bad-LSA or request errors and compare the LSDB on both sides. A software defect is possible, but first rule out ordinary topology and configuration problems. If the issue appeared after redistribution or area changes, inspect the external or summary LSAs introduced by that change.
Avoid clearing the OSPF process before capturing the stuck state. A reset can temporarily synchronize the databases and erase the specific request that identified the problematic LSA.
Read the LSDB according to area and LSA scope
Once neighbors are Full, missing routes become a database question. Determine which LSA should represent the destination in the local area. Intra-area topology is described by router and network LSAs. Inter-area reachability is represented through summary information from ABRs. External routes use external LSAs, while NSSA external routes begin as Type 7 inside the NSSA.
If the LSA exists in the origin area but not beyond the ABR, inspect summarization, filtering, area type, and ABR behavior. If a redistributed route is missing, check the ASBR and the redistribution policy. Do not jump directly to OSPF cost when the route information never crossed the expected boundary.
The LSDB should be examined on the router that performs the transformation. For example, an NSSA ABR is the useful place to verify Type 7 to Type 5 translation. A remote internal router only shows the result, not why the translation succeeded or failed.
When multiple areas are involved, compare the same prefix on an internal router, the ABR, and a router in the remote area. This three-point view shows where an intra-area route became an inter-area summary or where it disappeared. It is much faster than changing costs on the destination router when the ABR never advertised the route in the first place.
Check path preference when the LSA exists but the route does not
OSPF can know a destination in the LSDB without installing that path in the IP routing table. Another routing source may have a lower administrative distance, or the OSPF path may depend on an unreachable forwarding address or next hop. Compare the OSPF calculation with the final RIB decision.
Within OSPF, route type and cost affect preference. Intra-area, inter-area, and external routes are not all treated identically, and external type 1 and type 2 metrics have different calculation behavior. Identify the installed route source and metric rather than assuming the lowest visible OSPF cost always wins.
The broader OSPF routing model is useful here because BGP, static routes, and other IGPs can legitimately compete in the same RIB. The missing OSPF route may be a preference result rather than a protocol failure.
LSA age and sequence information can also reveal churn. Repeatedly refreshed or rapidly changing LSAs may point to an unstable link or flapping redistributed source. A stable adjacency can coexist with a noisy LSDB, so neighbor uptime alone is not enough for performance incidents. Monitor SPF frequency and LSA generation when CPU spikes or route convergence is unexpectedly frequent.
Troubleshoot OSPFv3 with address-family and link-local context
OSPFv3 uses different packet and addressing behavior from OSPFv2 even though the link-state concepts are similar. Verify the correct address family, interface activation, router ID, and link-local neighbor context. A dual-stack link can have healthy OSPFv2 adjacency while OSPFv3 is absent or misconfigured.
Check authentication and IPsec or platform-specific security mechanisms according to the software design being used. Do not assume an IPv4 authentication template applies unchanged. Also verify that IPv6 filtering does not block the protocol or required neighbor-discovery behavior.
When comparing routes, use the OSPFv3 LSDB and IPv6 routing table together. The same troubleshooting sequence applies: neighbor, database, calculation, installation, forwarding, but the commands and address-family context must match the protocol instance.
Create a small set of expected OSPF routes for each area and keep the associated route type, advertising router, and next hop in the operational baseline. During an outage, checking those representative prefixes quickly shows whether the problem is local to one area, an ABR boundary, external redistribution, or the final forwarding path.
After any repair, verify that the database is stable rather than merely synchronized once. Watch the neighbor for several dead intervals, confirm that LSA counters stop changing unexpectedly, and repeat representative traffic tests. Intermittent MTU, authentication, or physical errors can allow a brief Full state before the adjacency fails again, so a single successful command output is not enough evidence of recovery.
Keep neighbor-state history where monitoring supports it. A device that is Full at the moment an engineer logs in may have reset five times during the previous hour. Correlating adjacency changes with interface errors, CPU spikes, and maintenance events can reveal an intermittent root cause that a one-time snapshot would miss.
For recurring incidents, automate collection of neighbor state, interface parameters, LSDB summaries, and selected routes when an adjacency changes. Consistent evidence shortens diagnosis and reduces the temptation to clear the process before the failure state has been captured.
Finish with forwarding tests and preserve evidence before resets
Once the neighbor is Full and the route is installed, test actual forwarding. Verify the recursive next hop, CEF or forwarding entry, ACLs, MTU, and return path. An OSPF control plane can be completely healthy while user traffic fails elsewhere. Use ping and traceroute with deliberate source addresses so the test follows the same path as the affected application.
Preserve state before clearing neighbors or processes: neighbor output, interface OSPF parameters, LSDB entries, routing table, logs, and relevant packet captures. The general troubleshooting discipline of changing one variable at a time is especially important with OSPF because resets can hide transient database or timer problems.
After the fix, reproduce the original failure condition where safe. If the problem was a mismatched MTU, verify a large-packet test. If it was an area or authentication mismatch, validate both sides after the configuration is standardized. A stable Full neighbor is the first success criterion; correct routes and successful traffic are the final ones.