Replication is the continuous or frequent copying of data, or of whole systems such as virtual machines and databases, from a primary location to a second one, so that a current copy is ready if the primary fails. It is the engine behind most disaster recovery services and many high-availability designs. Replication can be synchronous, where both copies are written before the change is confirmed, or asynchronous, where the second copy trails slightly behind.
At a glance
- Replication keeps a second copy close to current, which supports fast recovery with little data loss.
- Synchronous replication keeps both copies in step but usually needs short distances and low-latency links; asynchronous replication tolerates distance but can lose the last few seconds or minutes of changes.
- It can work at the storage, hypervisor, database or application level, depending on the product.
- Replication copies bad changes too, including deletions, corruption and ransomware, so it doesn’t replace backup.
- It is a common foundation for disaster recovery as a service (DRaaS).
What problem it solves
Restoring from a nightly backup can take hours or days and loses everything since the last backup ran. For systems where that is unacceptable, such as order processing, patient records or core databases, organizations need a copy that is both current and ready to run. Replication provides that copy by sending changes to a second site as they happen or soon after. When the primary fails, systems can be brought up from the replica, often in minutes, with a much smaller data loss window than a daily backup. That directly supports a tight recovery point objective (RPO) and recovery time objective (RTO).
How it works
Initial copy, then changes. Replication starts with a full copy of the protected data at the target. After that, only changes are sent, which keeps bandwidth needs tied to how much data changes rather than total size.
Synchronous replication. Each write is confirmed to the application only after it is stored at both locations. The replica stays in step with the primary, but every write waits for the round trip, so this is usually limited to sites relatively close together with fast, reliable links. It is common for high availability (HA) within a campus or metro area.
Asynchronous replication. Writes are confirmed locally and sent to the target shortly after, in a stream or in frequent batches. This works over long distances and ordinary WAN links, at the cost of a small lag. If the primary fails, changes not yet sent are lost.
Where it runs. Storage arrays can replicate volumes, hypervisors and DR tools can replicate virtual machines, databases have their own replication features, and some applications replicate at the application layer. Each level has different trade-offs in what is protected and how recovery works.
Recovery points. Some replication tools keep a journal of recent changes so you can recover to a chosen moment rather than only to the latest state. Others pair replication with periodic snapshots. This matters when the latest replicated state is itself damaged.
Failover and back. In an outage, systems are started from the replica through failover; once the primary is fixed, data is synchronized back and service returns through failback.
When it matters for buyers
- When setting recovery targets. Replication is usually how short RPO and RTO targets become achievable.
- When choosing a DR service. DRaaS providers differ in replication method, how many recovery points they keep and how failover is triggered.
- When planning a second site or cloud region. Distance and connectivity decide whether synchronous replication is realistic.
- After something broke. An outage often reveals that a “replica” was out of date or never tested.
Questions to ask vendors
- Is replication synchronous or asynchronous, and what lag should we expect in normal operation?
- At what level does it run (storage, hypervisor, database or application), and what does that cover?
- How many recovery points do you keep, and how far back can we recover if the latest copy is corrupted or encrypted?
- What bandwidth will we need for the initial copy and ongoing changes?
- How are failover and failback performed, and how often can we test without affecting production?
- What alerts tell us if replication falls behind or stops?
How it differs from backup
Replication and backup are often confused because both create copies. Replication aims to keep one copy as current as possible, ready to take over quickly. Backup keeps a series of older, point-in-time copies so you can go back to a moment before something went wrong. Since replication passes changes on quickly, a deleted table or ransomware encryption can reach the replica too. Backup, especially isolated or immutable backup, gives you a clean point to return to. Most resilient designs use replication for speed and backup for history. Our disaster recovery as a service overview covers how providers combine them.
