FortiGate high availability is not just a pair of firewalls mounted next to each other. A resilient design must synchronize configuration and appropriate session state, detect device and path failures, select a primary unit predictably, preserve network reachability during transition, and make failures visible to operations. FortiGate Clustering Protocol (FGCP) provides the foundation for this on supported FortiGate deployments, and HA behavior is a practical part of the current FortiGate 7.6 administration skill set.
The central design question is which failures the cluster is expected to survive. A second appliance protects against one class of failure, but it does not automatically protect against a failed switch, a broken upstream route, asymmetric topology, or a configuration error synchronized to both members. Good HA engineering therefore begins with failure domains and traffic paths, then uses cluster settings to support that architecture.
Know what an FGCP cluster is protecting
An FGCP cluster is built from compatible FortiGate units that cooperate as one security system. Fortinet’s current FortiOS documentation requires cluster members to use compatible hardware and firmware, and the units share most configuration while retaining member-specific values such as host names and HA priorities. That consistency is necessary because a failover is useful only if the new primary can enforce the same policy and participate in the same network design.
HA does not remove the need for redundancy around the firewall. If both cluster members connect through one access switch, power source, or upstream router, that common dependency remains a single point of failure. Map the entire path from inside hosts through the cluster to external networks and identify which components have independent alternatives.
The broader availability lesson is similar to the reasoning behind high-availability targets: availability is an end-to-end property. A cluster can improve firewall resilience, but the business service remains unavailable if DNS, routing, switching, authentication, or the application itself has no comparable recovery design.
Choose active-passive or active-active deliberately
In an active-passive design, one unit handles traffic while the secondary is ready to assume the primary role. This is often easier to reason about because a single member is the active forwarding point. Active-active designs can distribute some processing responsibilities, but they do not mean every session is independently load-balanced across two unrelated firewalls.
The mode should reflect operational requirements, platform behavior, and the team’s ability to troubleshoot it. If the primary goal is deterministic failover with minimal complexity, active-passive is often attractive. If active-active is selected, administrators should understand which traffic or security processing can be distributed and what still depends on the primary member.
Cross-vendor explanations of active-active firewall failover can help illustrate the architectural trade-off, but FortiGate implementation details must be learned from FortiOS. Similar terminology does not imply identical election, session, or interface behavior.
Design heartbeat links as cluster infrastructure
Heartbeat interfaces carry the control traffic that lets cluster members discover one another, exchange status, synchronize configuration and state, and participate in primary selection. They are therefore part of the cluster control plane, not ordinary user interfaces. Dedicated and redundant heartbeat connectivity reduces the chance that a single cable or port failure causes the members to lose cluster visibility.
Heartbeat placement should avoid shared failure domains where practical. Two logical heartbeat links that ultimately traverse the same switch and power source are not as independent as two physically separate paths. On larger platforms, interface choice and bandwidth also matter because state synchronization and monitoring create ongoing traffic between members.
Operators should monitor heartbeat health before an incident occurs. A cluster can appear normal while already running with reduced heartbeat redundancy. If the remaining path later fails, troubleshooting becomes far more urgent. Treat a failed secondary heartbeat as a fault to repair, not as a harmless condition merely because production traffic is still flowing.
Understand primary selection and override behavior
FGCP uses election criteria to determine which member becomes primary. Device health, monitored interfaces, priority-related settings, and cluster history can influence the result depending on configuration. Administrators should understand their chosen behavior well enough to predict which unit will be primary after a reboot or failure.
Priority settings are often misunderstood as a simple permanent preference. In a stable cluster, avoiding unnecessary failback can be more valuable than forcing a preferred appliance to resume primary status immediately after recovery. Every extra transition creates another period of control-plane change and potential session impact.
If the organization requires a particular member to be primary—for example because of cabling, locality, or operational procedures—document why and test the resulting election behavior. If there is no such requirement, stability may be preferable to aggressive override. The correct setting is one that supports the physical design and recovery objective.
Monitor interfaces and upstream reachability
An appliance can remain powered on while losing a critical data interface. Interface monitoring lets the cluster treat selected link failures as health problems that may justify failover. The monitored set should represent paths whose loss makes the member unable to perform its intended role.
Do not monitor every interface reflexively. A noncritical lab or unused interface should not trigger a production failover. Conversely, monitoring only the obvious WAN port may miss the inside dependency that actually isolates applications. Select interfaces based on the end-to-end traffic architecture.
Link state also has limits. An Ethernet interface can remain up while the upstream network is unusable farther away. Routing protocols, SD-WAN health checks, or external monitoring may be needed to detect those failures. HA design is strongest when appliance health and network-path health complement one another rather than being treated as the same signal.
Plan session synchronization around application tolerance
Failover is most visible through session behavior. Without useful state synchronization, a new primary may need clients to establish fresh sessions after transition. FortiOS HA provides session synchronization capabilities, including session pickup options, to reduce disruption for supported traffic. The exact behavior depends on protocol, configuration, and the state that can be synchronized.
Application requirements should drive how much continuity matters. Short web requests may recover quickly through retries, while long-lived administrative sessions, voice traffic, transactions, or tunnels can make even a brief state loss noticeable. Document which traffic classes have strict continuity requirements and test them during planned failovers.
Session pickup is not a substitute for application resilience. Some applications bind state to a connection in ways the network cannot fully preserve, and upstream peers may react to topology changes. The useful question is not “did the cluster fail over?” but “did the business flow meet its recovery objective?”
Account for routing, switching, and MAC behavior
A successful primary transition must be reflected in the surrounding network. Switches and routers need to direct frames and packets toward the active member, dynamic routing neighbors may need to reconverge, and upstream devices may cache forwarding information. The firewall’s internal election can be fast while the end-to-end path still takes longer to recover.
Designers should understand how the cluster represents interfaces and addresses on the network, how adjacent devices learn the active forwarding location, and whether link aggregation or virtual-switch constructs introduce additional dependencies. Testing only from the firewall console can miss delays visible to clients.
Routing design also affects symmetry. Stateful firewalls need both directions of a session to be processed consistently. A failover topology that creates an unexpected return path can break traffic even though the new primary is otherwise healthy. Review routing on both sides of the cluster as part of every HA change.
Test failover with controlled failure scenarios
An HA configuration is incomplete until it has been tested. Planned tests can include loss of the primary unit, failure of a monitored interface, heartbeat degradation, reboot, software maintenance, and recovery of the original member. For each case, record election timing, packet loss, session behavior, routing convergence, logs, and alerts.
Testing should include applications, not only pings. A ping can survive while TLS sessions reset, a VPN renegotiates, or a database connection fails. Select representative flows and measure the behavior that users care about. Where possible, run the tests during commissioning and after significant network changes.
Build an observation checklist before the test so the team captures evidence rather than relying on impressions. Note which unit became primary, why the election occurred, whether both members synchronized afterward, and whether the recovered unit rejoined cleanly. Those records also become a reference for real incidents.
Configuration synchronization should be verified as a health signal in its own right. A secondary that is physically present but out of synchronization is not a trustworthy recovery target. After major policy, interface, routing, or certificate changes, confirm that members report the expected synchronized state and investigate persistent differences rather than waiting for a failover to expose them.
Maintenance procedures should distinguish graceful role changes from uncontrolled failures. When upgrading firmware or replacing hardware, administrators can plan which member is active, observe synchronization, and move traffic under controlled conditions. That produces much better evidence than treating every maintenance window as a surprise power-loss test.
HA and IPsec or dynamic routing can interact in subtle ways. Peers outside the cluster may need to renegotiate tunnels or routing adjacencies after a transition, and convergence time can exceed the cluster election itself. Include representative VPN and routing neighbors in failover testing instead of measuring only local interface recovery.
Split-brain risk also belongs in the design discussion. If cluster members lose the ability to communicate over heartbeat paths while both retain data connectivity, the architecture must prevent conflicting forwarding behavior. Redundant heartbeat links and well-understood cluster state reduce this risk; monitoring should alert on degraded heartbeat before the remaining control path is lost.
Finally, track recovery objectives quantitatively. Record packet-loss duration, application reconnect time, tunnel recovery, routing convergence, and time until the cluster returns to a fully synchronized redundant state. Those numbers turn “HA seems fast” into an engineering result that can be compared with the service’s actual availability requirement.
Firmware upgrades add version compatibility to the HA equation. Follow the supported upgrade path and verify cluster status before starting; an already-degraded cluster is a poor place to begin maintenance. During the upgrade, observe member versions, role changes, synchronization, and traffic continuity, then confirm both units return to the expected healthy state.
Hardware replacement should be rehearsed as a lifecycle task rather than improvised during failure. Keep records of licensing, interface mapping, HA credentials, management access, and the procedure for adding a replacement member without overwriting the active configuration. Spare capacity is only useful if the team can bring it into the cluster safely.
Alerting should distinguish a failover event from a loss of redundancy. A successful transition may restore service within seconds, but the environment is still exposed until the failed member is repaired and synchronization is restored. Operations should track both milestones: service recovered and redundancy recovered.
Finally, perform a post-failover review even when users saw no outage. Identify the trigger, confirm the cluster responded for the expected reason, and check for repeated flaps or degraded links. Silent success can still reveal a failing switch port or power component that will cause a larger problem if left unresolved.
Capacity must remain adequate on a single member after failure. A two-node cluster is not resilient if normal load requires both appliances to remain active at their combined maximum. Size CPU, memory, session, VPN, and inspection requirements so the surviving unit can carry the expected failover load without immediately becoming a second bottleneck.
Operational access should also survive the event. Confirm that administrators can reach the new primary through management networks, console paths, or out-of-band systems after a failover. Losing production traffic and management access at the same time turns a routine HA event into a much harder incident.
HA monitoring should include environmental differences between members. Power feeds, transceivers, cabling, and switch ports can fail independently even when the FortiGate configuration is synchronized. Document those physical dependencies so a repeated member failure is not mistakenly treated as a software problem.
Document normal cluster status after commissioning, including role, heartbeat links, synchronization indicators, and monitored interfaces. During an incident, comparing current output with a known-good baseline helps distinguish an actual HA fault from unfamiliar but normal status information.
Test architectural assumptions before incidents occur. A planned exercise can reveal whether monitored interfaces, heartbeat paths, routing neighbors, management access, and application sessions respond the way the design document predicts, giving the team time to correct gaps without outage pressure.
Troubleshoot HA as a state machine, not a mystery
When a cluster behaves unexpectedly, establish member state first: cluster membership, primary/secondary roles, heartbeat status, synchronization state, monitored-interface status, and recent election events. Then compare that evidence with the expected election logic. This is more reliable than rebooting members until the preferred one appears primary.
If failover succeeds but traffic does not, move outward. Check interface state, ARP or neighbor behavior, routing, session pickup, NAT, tunnels, and upstream/downstream reachability. If traffic works but sessions reset, focus on state synchronization and application tolerance. If the wrong member wins an election, focus on HA settings and monitored health signals.
High availability is therefore an operational discipline as much as a configuration feature. Candidates progressing through the NSE 4 FortiOS Administrator path should be able to explain not only how to form a cluster, but which failure a setting is meant to handle, what evidence proves the cluster is healthy, and what the surrounding network must do for failover to become real service continuity.