A load balancer is a device, piece of software or cloud service that sits in front of a group of servers and distributes incoming requests among them. By spreading the work, it helps keep individual servers from being overwhelmed, and by checking each server’s health, it stops sending traffic to servers that fail or are taken down for maintenance. Users see one address; behind it, the load balancer decides which server handles each request. Load balancers are a standard part of websites, applications and APIs that need to scale or stay available.
At a glance
- A load balancer presents one address to users and spreads requests across multiple servers behind it.
- Health checks let it stop sending traffic to servers that fail, reducing the impact of a single failure.
- It is available as hardware appliances, software and managed cloud services.
- Layer 4 load balancers route by network address and port; Layer 7 load balancers read application requests such as web URLs.
- The load balancer itself must be redundant, or it becomes the new single point of failure.
What problem it solves
An application on one server has two weaknesses: it can only handle as much traffic as that server can, and when the server fails or needs patching, the application is down. Adding more servers solves the capacity problem only if something divides the traffic between them, and solves the availability problem only if something notices when one fails.
That is the load balancer’s job. It lets a team add servers to handle growth, remove servers for maintenance without an outage, and limit the impact of a crash to the requests that server was handling. It is a building block of high availability (HA) for web and application services.
How it works
Distribution methods. Common methods include round robin (each server in turn), least connections (the least busy server) and hashing (the same client goes to the same server). Weights can send more traffic to larger servers.
Health checks. The load balancer regularly tests each server, for example by requesting a web page. Servers that fail stop receiving new traffic until they pass again, which also supports failover between groups of servers.
Layer 4 versus Layer 7. Network-level (Layer 4) balancing forwards connections by IP address and port. Application-level (Layer 7) balancing reads requests, so it can route by URL path or hostname, terminate TLS encryption, add headers and, in many products, apply security rules.
Session persistence. Some applications need a user to stay on the same server during a session; load balancers can pin sessions using cookies or source addresses, at some cost to even distribution.
Global load balancing. Some services balance traffic across sites, regions or availability zones, often using DNS or anycast routing, so users reach a nearby healthy location.
Delivery options. On-premises, organizations use hardware appliances or software on servers or virtual machines. Public clouds offer managed load balancers that can scale automatically and are usually placed inside a virtual private cloud (VPC).
For edge delivery options that work with load balancing, see our cloud content delivery network solution page.
When it matters for buyers
- When an application outgrows one server. Adding servers behind a load balancer is the usual next step.
- When planned downtime is no longer acceptable. Rolling updates through a load balancer let teams patch servers one at a time.
- When moving to the cloud. Managed cloud load balancers replace appliances but are priced differently, often per hour plus data or connections processed.
- When renewing appliance support. Hardware load balancer refreshes are expensive; compare against software and cloud services.
- When adding security. Many Layer 7 load balancers can integrate with or include web application and API protection features.
Questions to ask vendors
- Do you support Layer 4, Layer 7 or both, and which protocols?
- How is the load balancer itself made redundant, and what happens if one instance or zone fails?
- How are health checks configured, and how quickly is a failed server removed?
- How is the service priced (per hour, per connection, per gigabyte or per rule), and what does a typical month cost at our traffic levels?
- Can it terminate TLS, and how are certificates managed and renewed?
- Does it support global or multi-region balancing?
- Which security features are included, and which are separately licensed?
How it differs from a CDN
A load balancer distributes requests among your own servers, usually in one site, region or cloud. A content delivery network (CDN) is a distributed network of provider-run servers around the world that cache and serve content close to users, reducing the load on your servers and the distance data travels. They often work together: the CDN handles requests at the edge, and requests it cannot answer from cache pass to a load balancer in front of your origin servers. Some CDN providers also offer global load balancing between origin sites.
