Failover is the process of switching from a primary system, connection or site to a backup when the primary fails or becomes unavailable. It applies at many levels: from one internet circuit to another, from one server or firewall to its partner, or from one data center to a recovery site. The goal is to keep service running, or restore it quickly, with as little interruption as the design allows.
At a glance
- Failover needs two things: a backup to switch to and a way to detect failure and move traffic or workloads to it.
- It can be automatic, triggered by health checks, or manual, triggered by a person.
- Switching takes time; depending on the design, users may see anything from a short pause to dropped calls and sessions.
- A backup only helps if it does not share the failure: the same provider, conduit, power source or software fault can take out both.
- Untested failover is a common reason backups fail to take over in a real outage.
What problem it solves
Single points of failure are everywhere in IT: one internet circuit into an office, one firewall at the edge, one database server, one data center. When that component fails, everything that depends on it stops. Failover reduces the impact by having a standby ready to take over.
For a buyer, failover turns an outage that could last hours or days, waiting for a repair, into one that lasts seconds or minutes. It is how a business meets uptime expectations from customers, partners and its own staff, and it often features in continuity plans, insurance questionnaires and customer security reviews.
How it works
Detection. Something must notice the failure. Network devices track link status and send test traffic to confirm a path really works; servers and clusters exchange heartbeats; monitoring tools check application health. Detection settings balance speed against false alarms that cause unnecessary switching.
Switching. Once a failure is detected, traffic or workloads move to the backup. For internet circuits, a router or SD-WAN device redirects traffic to the second link. Sites with their own IP addresses can use Border Gateway Protocol (BGP) to announce them over both connections, so inbound traffic can find the surviving path. For servers, a standby takes over the service address or DNS is updated to point elsewhere.
State. Some failovers preserve sessions and data; others do not. Firewalls and clusters may synchronize state with their partners so existing connections survive. When the path or public address changes, many connections must be re-established.
Backup options for connectivity. Common pairings include fiber with cable broadband, fixed wireless or a cellular connection. Cellular failover is one widely used method because it rarely shares a physical path with wired circuits.
Failback. When the primary is healthy again, service returns to it, automatically or at a planned time.
Our SD-WAN page covers how multi-link designs handle failover between circuits.
When it matters for buyers
- After an outage. If a single failure stopped the business, failover is usually the first design change to consider.
- Opening a site that depends on the internet. Plan the backup connection with the primary, including its path into the building.
- Moving phones and core apps to the cloud. The internet link becomes critical, so its failover becomes critical too.
- Answering insurers, auditors or customers. Many ask how you would keep operating if a key system or site failed.
- Reviewing an SLA. An SLA credit compensates for downtime; failover is what limits it.
Questions to ask vendors
- What triggers failover, how is failure detected, and how long does detection and switching typically take?
- Which traffic and sessions survive a failover, and which must reconnect?
- Does the backup share any equipment, provider network, building entry or power with the primary?
- Is failback automatic or manual, and can we control when it happens?
- How do we test failover, and how often do you recommend it?
- Who is alerted when failover happens, and how do we know when we are running on the backup?
- If you manage the service, is failover monitoring and testing included?
How it differs from high availability
High availability (HA) is a design goal: keeping a service running for as large a share of time as practical, through redundancy, monitoring and careful engineering. Failover is one of the mechanisms used to get there, the act of moving to the backup when something fails. A highly available system usually includes failover, but it also includes things like removing single points of failure, maintenance without downtime and capacity planning. A system can have failover and still not be highly available, for example if switching is manual and slow or if both sides share a hidden dependency.
