N+1 redundancy is a design approach in which a system has one more component than it needs to carry its full load. “N” is the number of units required, such as uninterruptible power supply (UPS) modules, cooling units, generators or servers in a cluster, and the “+1” is a spare. If one unit fails or is taken out for maintenance, the remaining units can still carry the load. It is one of the most common ways data centers and IT systems describe their resilience.
At a glance
- N is what the load requires; N+1 adds one spare, designed so a single failure or maintenance event does not reduce capacity below what is needed.
- It applies to many kinds of components: power, cooling, generators, network devices and servers.
- N+1 protects against one component failing within a redundant group; it does not by itself protect against a failure of the shared distribution path.
- Higher levels include N+2, 2N (two complete systems) and 2N+1.
What problem it solves
Every component eventually fails or needs maintenance. In a system with exactly N components, losing one means losing capacity: a cooling plant with no spare may not keep up, and a UPS system with no spare may drop the load or have to switch it to unprotected bypass. Maintenance then has to wait for a shutdown window.
N+1 removes that dependence on every unit working. With one spare in the group, the system can lose a unit and keep running at full capacity, and technicians can service units one at a time. It is a relatively economical way to remove single-component failures as a cause of downtime, which is why it appears in data center designs, network equipment and server clusters alike.
How it works
Sizing N. Designers first determine how many units are needed for the full design load. That figure depends on the load assumed, so N for a half-full facility may differ from N at full capacity.
Adding the spare. One additional unit is installed and runs alongside the others or stands by to take over. In active designs, all units share the load at lower utilization; in standby designs, the spare starts when needed. A cold spare sitting on a shelf is a related but slower approach, since someone must install it.
Higher levels. N+2 adds two spares. 2N provides two fully independent systems, each able to carry the whole load, typically feeding separate power or cooling paths. 2N+1 adds a spare on top of that. Each step costs more and protects against a wider range of failures.
Distribution paths. N+1 is usually applied to components such as UPS modules or chillers. If all those components feed a single path to the equipment, a fault in that path can still cause an outage. This is why tier classifications look at distribution paths as well as component counts.
In IT systems. The same idea applies to clusters. A server or hyperconverged cluster sized N+1 keeps enough spare capacity to run all workloads if one node fails, with failover moving the affected workloads to the remaining nodes.
For help evaluating the redundancy of a facility or platform, see our Colocation solution page.
When it matters for buyers
- Choosing a data center. Ask what redundancy each part of the power and cooling chain has, not just the headline claim.
- Reviewing a failure. After an outage, check whether a single point of failure sat outside the redundant group.
- Sizing a cluster. Make sure a hyperconverged or virtualization cluster can carry all workloads with one node down.
- Growing into a facility. As load grows, an N+1 system can quietly become N if spare capacity is used up.
Questions to ask vendors
- What is the redundancy level (N+1, 2N or other) for UPS, generators, cooling and power distribution?
- Is N calculated at today’s load or at the facility’s full design load?
- Do the redundant components share a single distribution path, switchgear or control system?
- How do you prevent spare capacity from being used up as more customers or workloads are added?
- How is maintenance performed on redundant equipment, and does it ever reduce redundancy below N+1?
- For clusters, how much capacity is reserved to absorb a node failure?
How it differs from high availability (HA)
N+1 is one specific technique: adding a spare component to a group. High availability is the broader goal of keeping a service running through failures, measured by its uptime, and it combines many techniques, including redundancy at several layers, failover, monitoring and operational practice. An N+1 power plant contributes to high availability, but a highly available service also needs its network, software and data protected against failure.
