Failback is the process of moving systems, data and users from a backup environment, such as a disaster recovery site or cloud recovery platform, back to the primary environment after it has been repaired or rebuilt. It is the second half of a recovery: failover gets you running again during an outage, and failback returns you to normal operations without losing the work done in the meantime. Failback is often harder than failover, and less often tested.
At a glance
- Failback reverses a failover once the primary site or system is healthy again.
- Data created or changed while running on the backup has to be copied back before switching over, or it is lost.
- For servers and applications it is usually a planned, scheduled event; for simple network links it may happen automatically.
- It often involves a short interruption, which is why timing and a written plan matter.
- Many disaster recovery tests cover failover only, leaving failback unproven.
What problem it solves
After a disaster, the urgent focus is getting systems running again somewhere. But a recovery site is rarely meant to be permanent: it may cost more per day, run at reduced capacity, sit in a provider’s environment with time limits, or lack integrations your normal systems rely on. At some point you need to move back.
Doing that badly can cause a second outage or lose data. If the business has been processing orders, updating records and receiving email on the recovery site for days, all of that has to come back with you. Failback provides the method: synchronize changes back to the primary, verify them, then switch users and traffic in a controlled way.
How it works
Restore the primary. The original site or system is repaired, rebuilt or replaced, and checked so the same failure or, after an attack, the same compromise isn’t brought back.
Reverse replication. Data changed on the recovery site is copied back to the primary. Many tools set up replication in the reverse direction, sending an initial bulk copy and then ongoing changes, so the primary catches up while the business keeps running on the recovery site.
Plan the cutover. A maintenance window is chosen. Applications are stopped or set to read-only on the recovery side, final changes are synchronized, and systems are started on the primary.
Switch users and traffic. Network routes, DNS records, VPNs and integrations are pointed back at the primary. The steps should be written in a runbook so they are done in the right order.
Verify and clean up. Teams confirm applications work and data is complete, then return the recovery site to standby and re-establish normal replication toward it.
Simple cases. For network failover, such as a backup internet link, failback can be automatic when the primary link is stable again, often with a delay so a flapping connection doesn’t cause repeated switches.
When it matters for buyers
- When choosing a disaster recovery service. Failback support, reverse replication and the time limits for running in the provider’s environment differ among disaster recovery as a service (DRaaS) providers.
- When setting recovery targets. Your recovery time objective (RTO) covers getting back online; failback is the separate work of returning to normal, and its duration and data risk need their own planning.
- When testing your plan. A test that stops after failover leaves the hardest part unverified.
- After something broke. If you are already running on a recovery site, failback planning starts now.
Questions to ask vendors
- Is failback included in the service, or billed as extra hours or a separate project?
- How do you replicate changes back to our primary, and how long would that take for our data volume?
- How long can we run in your recovery environment, and what does it cost per day after any included period?
- What downtime should we expect during failback for each system?
- Do your DR tests include a full failback, and can we see results from a recent one?
- Who performs each step, and is there a written failback runbook for our environment?
How it differs from failover
Failover and failback are two directions of the same move. Failover is triggered by a failure and focuses on speed: get service running on the backup with as little interruption and data loss as possible. Failback is usually chosen, not forced, and focuses on control: bring the primary back up to date, then switch at a time that suits the business. Failover often has the most attention and automation; failback more often relies on manual steps, which is why it deserves its own planning in your business continuity and disaster recovery (BCDR) program. Our disaster recovery as a service overview covers how providers handle both directions.
