What Is a Load Balancer and How It Works
SkillVeris Team
Cloud & Security Team

A load balancer distributes incoming network traffic across multiple servers so no single server is overwhelmed.
In this guide, you'll learn:
- It improves availability and scalability — if one server fails, traffic is routed to the healthy ones automatically.
- Health checks let the balancer detect and stop sending traffic to servers that are down or unresponsive.
- Algorithms like round robin, least connections, and IP hash decide which server handles each request.
- Layer 4 balancers route by IP and port; Layer 7 balancers understand HTTP and can route by URL, header, or cookie.
1What Is a Load Balancer?
A load balancer is a component that sits in front of a group of servers and distributes incoming requests across them. Instead of every user hitting one machine, the balancer spreads the load so each server handles a manageable share, keeping the application fast even under heavy traffic.
Think of it as the host at a busy restaurant directing arriving guests to whichever tables have free waiters. No single waiter gets swamped while others stand idle, and if one waiter leaves, the host simply stops seating people in that section. The result is a system that stays responsive and keeps working even when parts of it fail.
2Why Use a Load Balancer?
A load balancer solves two problems at once: handling more traffic than one server can, and staying online when a server dies. These are the core requirements of any serious web application.
- Scalability: add more servers behind the balancer to handle more users.
- High availability: if a server fails, traffic shifts to the others with no downtime.
- Performance: spreading load prevents any one server from becoming a bottleneck.
- Maintenance without downtime: take a server out of rotation, update it, and return it.
- Flexibility: route different requests to different server pools based on rules.
🔑Key Idea
A load balancer is what makes horizontal scaling possible — growing by adding many modest servers rather than buying one ever-bigger machine, which is cheaper and far more resilient.
3How a Load Balancer Works
When a request arrives, the load balancer chooses a backend server according to its configured algorithm and forwards the request there. The chosen server processes it and returns a response, which the balancer relays back to the client. To the user, it looks like a single fast, reliable server.
Crucially, the balancer constantly runs health checks — small periodic probes to each server. If a server stops responding, the balancer marks it unhealthy and stops routing to it until it recovers. This automatic detection is what turns a pool of servers into a system that survives individual failures.
💡Pro Tip
Point health checks at a dedicated endpoint like /healthz that verifies the app can reach its database and dependencies — not just that the web server is up. A server that returns 200 but can't reach its database should be pulled from rotation.
4Load Balancing Algorithms
The algorithm decides which server gets each request. Different algorithms suit different workloads, and picking the right one matters when requests vary in cost.
- Round robin: hand requests to each server in turn — simple and even when servers are identical.
- Least connections: send the next request to the server with the fewest active connections — better for long-lived requests.
- Weighted round robin: give more powerful servers a larger share of traffic.
- IP hash: route a given client's IP consistently to the same server — useful for session affinity.
- Least response time: favor the server currently answering fastest.
Sticky Sessions
Some applications store session state on individual servers, so a user must keep hitting the same one. Session affinity (sticky sessions) pins a client to a server. It works, but storing session state externally — in a shared cache or database — is more scalable because any server can then serve any user.
5Layer 4 vs Layer 7 Load Balancing
Load balancers operate at different layers of the network stack, and the distinction affects how smart their routing can be.
Layer 4 (Transport)
A Layer 4 balancer routes based on IP address and TCP/UDP port without inspecting the actual content. It is extremely fast and efficient because it doesn't read the request, but it can't make decisions based on what the request contains.
Layer 7 (Application)
A Layer 7 balancer understands HTTP. It can route based on the URL path, hostname, headers, or cookies — sending /api requests to one pool and /images to another, for example. It also enables features like SSL termination and content-based routing, at a modest performance cost.
6Types and Common Tools
Load balancers come as hardware appliances, software you run yourself, or managed cloud services. Most teams today use software or cloud options.
- Cloud managed: AWS Elastic Load Balancing, Azure Load Balancer, Google Cloud Load Balancing.
- Software: NGINX and HAProxy are widely used, fast, and free.
- Ingress controllers: in Kubernetes, an ingress controller acts as a Layer 7 balancer for cluster services.
- DNS load balancing: distributes traffic across regions by returning different IPs.
- Global vs regional: global balancers spread users across data centers worldwide.
7Common Mistakes to Avoid
Load balancers are reliable, but a few configuration errors undermine their benefits. Watch for these.
- Weak health checks that only ping the port, missing apps that are up but broken.
- Making the load balancer itself a single point of failure — use a redundant, managed one.
- Relying on sticky sessions instead of externalizing session state, which limits scaling.
- Ignoring connection draining, so in-flight requests are cut off when a server is removed.
- Forgetting to distribute across availability zones, leaving you exposed to a zone outage.
⚠️Watch Out
A load balancer with all its backends in one availability zone still fails entirely if that zone goes down. Spread backends across zones so the balancer has healthy targets to route to.
8Key Takeaways
The essentials of load balancing distill into a few core ideas.
- A load balancer distributes traffic across multiple servers for speed and availability.
- Health checks let it route around failed servers automatically.
- Algorithms like round robin and least connections decide who handles each request.
- Layer 4 routes by IP and port; Layer 7 understands HTTP and routes by content.
- It is the foundation of horizontal scaling — many servers instead of one big one.
9Frequently Asked Questions
Q: What is the difference between Layer 4 and Layer 7 load balancing? A: Layer 4 routes by IP address and port without reading the request, making it very fast. Layer 7 understands HTTP and can route by URL, header, or cookie, enabling smarter, content-based decisions at a small performance cost.
Q: Does a load balancer prevent downtime? A: It greatly reduces it by rerouting traffic away from failed servers, but only if it is itself redundant and its backends span multiple availability zones. A single load balancer in one zone can still be a point of failure.
Q: What are sticky sessions? A: Sticky sessions, or session affinity, pin a given client to the same backend server so server-stored session data stays available. Storing session state externally in a shared cache is more scalable because then any server can serve any user.
Q: Do I need a load balancer for a small site? A: Not until you run more than one server or need high availability. A single small site can run on one server, but as soon as you want zero-downtime deploys or to handle more traffic, a load balancer becomes essential.
Related Reading
Get The Print Version
Download a PDF of this article for offline reading.
About the Publisher
SkillVeris Team
Cloud & Security Team
Our cloud and security experts break down complex infrastructure topics into practical, beginner-friendly guides.
View all postsRelated Posts
Never miss an update
Get the latest tutorials and guides delivered to your inbox.
No spam. Unsubscribe anytime.