Quick summary
A load balancer distributes incoming requests across multiple servers. A health check measures whether a server can actually accept traffic, while sticky sessions try to keep a user on the same server. A reliable design considers all three together with session storage, TLS termination and failure handling.
- A health check should verify that the application is ready for traffic, not only that a port is open.
- Sticky sessions are a compatibility measure; shared session storage usually allows more flexible scaling.
- TLS termination determines where the encrypted client connection is opened and managed.
- Enabling a health endpoint and restricting safe access to it are separate tasks.
When you consider moving to multiple web servers, the first expectation is usually simple: distribute incoming traffic between them. But users being logged out, shopping carts appearing empty, or requests continuing to reach a failed server show that traffic distribution alone is not enough.
That is why the answer to what is a load balancer health check involves more than a single device feature. The load balancer selects requests, health checks identify suitable targets, and sticky sessions affect how session state behaves across servers. You can make better decisions by considering these parts together.
What is a load balancer and what problem does it solve?
A load balancer is a component that directs network connections or HTTP requests from clients to multiple backend servers. A backend is the web server, application server or service that actually processes the request. The user sees one domain name, while requests are distributed among suitable targets in the server pool.
In a single-server setup, hardware failure, maintenance or exhausted processing capacity affects the site directly. A load balancer can direct requests to other targets when one server is unavailable. It can also make it easier to add a server to the pool, introduce traffic gradually or remove a target temporarily during maintenance.
This design does not guarantee a faster site by itself. The load balancer, network connection, database, file storage or application code can still become a bottleneck. Before adding servers, review request types, session data, workload and the need for shared files.
The difference between L4 and L7 distribution
Layer 4 (L4) distribution makes decisions using network information such as TCP or UDP connections. It generally uses less application information and works at the protocol level. Layer 7 (L7) distribution can evaluate application data such as the HTTP method, host name, URL path or cookie.
Sending static files to one target group and API requests to another can be done with L7 rules. Passing a TLS connection through based only on connection information, without opening it on the load balancer, is closer to an L4 or TLS passthrough approach. The choice depends on your routing needs and where you want to manage certificates.
DNS can sit in front of the load balancer, with web servers behind it. DNS resolves a domain name to an IP address; it does not by itself check whether every HTTP request will reach a healthy target. Do not treat DNS routing and load balancer health checks as the same task.
What is a load balancer health check?
A health check is a recurring test that the load balancer sends to backend servers. When the test succeeds, the target remains in the pool. When failures pass a configured threshold, the target is temporarily removed. Normal user requests can then avoid a server that cannot respond or is returning incorrect responses.
A basic TCP check only shows that a port accepts connections. An HTTP check sends a request to a particular URL and can verify the status code and, when necessary, response content. The web server may appear to be running while PHP-FPM, the database or another critical application dependency is broken. Relying only on a TCP check can therefore mark a target healthy when the application is not ready.
What should a health endpoint check?
A health endpoint should respond quickly and predictably to show whether the application is ready to receive traffic. A path such as /healthz or /ready can be used. These names are examples; choose a path that fits your routing structure and operational policy.
It is useful to separate two types of checks. Liveness indicates that a process is running, while readiness indicates that the dependencies required to serve requests are available. Most websites need a readiness-like check for load balancing. Avoid running heavy report queries, full-page generation or long-running calls to external services in every health check, because the check can create additional load.
If you need to check the database, use a small, read-only query with a timeout. Decide in advance how the application should treat a temporary problem with an external dependency such as a payment service. Adding every external service to the health check can cause all servers to be removed from the pool during a short dependency outage.
Caution
Enabling a health endpoint is not the same as defining secure access to it from outside. Even if the application serves the path, allow access through the web server or load balancer only from expected networks. Do not put credentials in the response body, and return status information without unnecessary operational detail.
Why failure thresholds matter
Removing a server after one failed request can cause unnecessary switching during a temporary network delay. Waiting for too many failures can send traffic to a broken target for longer. Evaluate the interval, timeout, unhealthy threshold and healthy threshold together.
For example, checks might run at five-second intervals with a two-second timeout, but this is only an illustrative example. Set real values using your network latency, normal application response time and tolerance for failure. After changing them, observe whether the target actually leaves and rejoins the pool and how user requests behave.
Diagnose the check before changing application settings
If a server fails its health check, first test whether the load balancer can reach the check address on that target. Then call the same URL locally on the target with the correct Host header. This helps separate network, web-server routing and application-layer problems.
curl -i --max-time 3 -H "Host: example.test" http://127.0.0.1/healthzThis command is for observation only; replace the domain and path with those used in your setup. The prerequisites are curl on the target server, a local web server listening for HTTP requests and a defined health-check route. The option is supported by current curl releases; confirm the installed version with curl --version if your environment is unusual. If the response does not have the expected status code, fix the health-check route or the relevant web-server location rule first. Do not change unrelated application settings at the same time.
After the correction, verify three points: does the target return the expected status code, does the load balancer mark it healthy, and does a real user request receive the correct response? If the change affects data or routing rules, take a configuration backup first and keep the steps for restoring the previous file ready.
What are sticky sessions and when are they needed?
A sticky session is a method of directing the same client to the same backend server for subsequent requests whenever possible. It is also called session affinity. A load balancer can provide it with a cookie created by the load balancer, an application cookie, the source IP address or another routing key.
The need usually comes from session data being stored only on the server that handled the first request. A user logs in and the session is written to a local file on Server A. If the next request goes to Server B, that server may not find the session. The user may be sent back to the login page or see an empty cart.
Sticky sessions can reduce this problem quickly, but they do not remove the underlying data-design issue. Affinity can cause some servers to carry more sessions than others. If a server fails, users tied to it can still lose their sessions. Automatic scaling, maintenance and load distribution also become less flexible.
Cookie-based and IP-based affinity
With a cookie-based method, the load balancer or application gives the client a cookie that identifies the preferred target. This reduces the problem of many users behind the same corporate proxy or NAT appearing as one source IP. Affinity can still break when the cookie is deleted, expires or is handled differently by the client.
With an IP-based method, the source IP is used to select a target. Mobile networks, corporate proxies and NAT can make one IP represent many people. A user can also move to another server when their IP changes. Do not treat IP affinity alone as a reliable solution for session security or data consistency.
A more scalable approach is to move session data to a shared store that all application servers can access. Redis, Memcached or a shared database can serve this purpose; the right choice depends on data persistence, access time and tolerance for loss. Define access permissions and expiration policies along with the storage location.
In WordPress, WooCommerce or a custom PHP application, verify the session and cache behavior first if file-based sessions remain tied to one server. Before switching to shared storage, plan how existing sessions will be ended and how connection details will be recovered. Back up the relevant files and data before changing the configuration, and keep the previous session configuration available for rollback.
For a related comparison of storage approaches, see PHP session and cache storage options. Review the application documentation and your current hosting limits before choosing between files, Redis and Memcached.
Example scenario
Consider a two-server WooCommerce site. Assume that cart sessions are stored in local files on each server while both servers use the same database. A user reaches Server A, adds a product and is sent to Server B on the next request. Server B may not find the session file created on A. Sticky sessions can be a temporary fix; a more durable approach is to move session data to shared storage and then retest whether affinity is still needed.
TLS termination and request-routing decisions
TLS encrypts the connection between a client and a service. TLS termination means that the encrypted connection is opened at a particular point and the HTTP request is passed to the next component. If termination takes place on the load balancer, certificates are managed there; the backend connection can use HTTP or be encrypted again with HTTPS.
Terminating TLS on the load balancer can centralize certificate renewal and L7 routing rules. However, you must assess the trust of the network between the load balancer and backend servers, internal network policy and sensitive-data requirements. If you also use TLS internally, plan certificate validation, name resolution and the trust chain.
Verify that the application correctly handles forwarded headers for the client IP, protocol and host. Do not allow untrusted clients to spoof these headers. Configure the web server and application to accept forwarded information only from the trusted proxy layer. After changing this configuration, verify redirects, absolute URLs, access logs and any IP-based security rules.
How should web and background workloads be separated?
A load balancer distributes web requests, but every backend task does not consume resources in the same way. PHP-FPM web workers process incoming HTTP requests. CLI cron jobs, queue consumers and scheduled commands do not consume a PHP-FPM worker unless they make an HTTP call. Mixing these worker types in one capacity calculation can lead to an incorrect diagnosis.
For example, a health check can respond quickly while long-running queue jobs consume CPU, memory or database connections. In that case the load balancer is not necessarily faulty; web and background workloads may be competing for resources on the same server. A separate process pool can help you treat web requests and queue workers with separate resource policies. The existing guidance on separate PHP-FPM process pools is a relevant next reference for this design.
During diagnosis, inspect PHP-FPM status, web-server access logs, application error logs and queue-worker status separately. Application processes must be ready for the health check, but a queue worker running or stopped should not by itself determine whether a web target is healthy.
Testing and observation before production traffic
Before placing two servers into production, write down which target will handle which requests. Check whether static files are identical on both servers, where uploads are stored, how sessions are shared and how database connections are managed. If a file created on one server is not available on the other, distributing traffic alone can produce inconsistent responses.
Then run a controlled capacity test. A test with one virtual user does not show real concurrency or queue behavior. Use a workload that reflects login, product browsing, cart and payment flows. The guidance on load testing with k6, JMeter and Locust can help you organize those scenarios.
Do not monitor only average response time. Review the error rate, 95th or 99th percentile response time, health-check failures, connection pools, CPU, memory and database latency together. Derive numerical thresholds from your application’s normal operating measurements; there is no universal value.
Remove a failed target in a controlled way
Before disabling a target for maintenance, stop new connections or use a draining feature if one is available. Draining allows existing connections to finish while new requests go to other targets. If you use sticky sessions, also test what happens to existing cookies and long-lived connections.
Next, run a single-target failure scenario. Observe the target failing its health check, leaving the pool and the remaining server continuing to answer requests. When bringing the target back, do not immediately send it full traffic. First check the health result, application logs and resource usage.
Frequently Asked Questions
Can a load balancer use a health check without an application endpoint?
Yes. A TCP check can test whether a port accepts connections, but it cannot confirm that the application is ready or that critical dependencies work. An application-level readiness check gives you more useful information when the platform supports it.
Should every backend server use the same health-check path?
They should use a consistent policy when they serve the same application role. Different roles may need different endpoints, but the load balancer should know what each response means before targets are placed in the same pool.
Can a health check cause an outage?
It can contribute to one if it is expensive, has an unsuitable timeout or depends on a failing external service. Keep it lightweight, define thresholds carefully and test the removal and recovery behavior before production use.
Does failover preserve existing WebSocket connections?
Usually not. A failed backend connection generally has to be re-established by the client or application. Test reconnect behavior separately from ordinary short-lived HTTP requests, especially when sessions or unsaved state are involved.
Actionable checklist
- Document the role of each backend, shared-data requirement and traffic path.
- Separate a TCP check from an application-readiness check.
- Confirm that the health-check path is fast, inexpensive and limited to necessary dependencies.
- Allow access to the check endpoint only from trusted load-balancer sources.
- Verify whether session data is stored in local files or shared storage.
- If you use sticky sessions, document session-loss and load-imbalance risks during server failure.
- Decide where TLS certificates will be terminated and how the backend connection will be protected.
- Monitor PHP-FPM web workers separately from CLI cron and queue workers.
- Test draining, health-check recovery and restoration of the previous configuration.
- Repeat capacity and failure tests with real user flows before production migration.
Your next step can be to prepare a read-only readiness endpoint on one existing server and verify it with a local request. Then run a single-target load-balancer test. Do not move all traffic to the new architecture until health status, session behavior and the recovery procedure are clear.





