The Midnight Single Point of Failure (SPOF) Nightmare
At exactly 11:00 PM on Black Friday 2023, the power supply on our Nginx reverse proxy server blew up. Behind it, a cluster of four backend nodes and a database cluster were running smoothly, with CPU load under 30%. Yet, our customers were instantly greeted by Connection timed out errors. The reason was simple: our domain pointed directly to the single public IP of that proxy. Once the gateway went down, the entire system was paralyzed.
I rushed to the Cloudflare dashboard to switch the A record to the standby server’s IP. Even though Cloudflare propagates record changes in mere seconds, visitors on local ISPs were stuck with cached responses for nearly 10 minutes. Around 350 orders vanished into thin air. Over 120 million VND in revenue was lost. That was the steep price paid for an architecture with a single point of failure (SPOF).
Why DNS Failover Can’t Save You in Time
Many people assume: “Why overcomplicate things? Just use DNS round-robin or run a cron job to update records via the Cloudflare/Route53 API!”. But in real-world operations, this approach reveals three fatal flaws:
- Stubborn DNS caching by clients and ISPs: Telecom providers like VNPT, Viettel, and FPT, as well as client browsers, frequently ignore short TTLs (such as 60s). Many sessions remain stuck hitting the dead IP for 15 to 30 minutes.
- Blind to service health: DNS only resolves IP addresses. It has no idea whether Nginx is still alive or just crashed due to an out-of-memory error (OOM Killer).
- Excessive detection latency: A health-check script usually requires 3–5 failed pings/curls before triggering an IP change. Factor in API call latency, and the actual downtime often stretches past 3–5 minutes.
The root problem lies at the network layer. We cannot delegate high availability to DNS at the external application layer. Instead, use a Virtual IP (VIP) within the local network. At this layer, servers negotiate and hand over control in just a few milliseconds.
Three High Availability Solutions for Gateways and How to Choose
In Linux environments, systems engineers typically weigh three main options:
- Cloud Load Balancers (AWS ALB, GCP LB): Hands-off and reliable, with provider SLAs reaching 99.99%. The downside is skyrocketing costs under heavy throughput (hundreds of Mbps or higher). Furthermore, this approach isn’t viable on on-premises infrastructure, dedicated servers at colocation facilities (like Viettel IDC or CMC), or self-hosted Proxmox/KVM clusters.
- Pacemaker + Corosync: An enterprise-grade clustering stack with robust split-brain fencing and complex resource management. The trade-off is high configuration complexity. In a two-node cluster, quorum loss is a constant risk unless you configure a qdevice properly.
- Keepalived (VRRP Protocol): Lightweight, consuming under 20MB of RAM, and configured entirely within a single file. Keepalived continuously broadcasts heartbeats across the local network. When the primary node fails, the standby node claims the Virtual IP onto its network interface within 1–2 seconds.
Standard Keepalived + VIP Deployment on Ubuntu Server
For most mid-to-large web setups running Nginx or HAProxy, Keepalived strikes the ideal balance between stability and operational overhead. Much like designing ingress and load balancing in a multi-node Kubernetes cluster, a solid entry-point strategy protects upstream applications. Below is a standard Master–Backup architecture you can deploy immediately.
Assumed Network Parameters
- Node 01 (Master): IP
192.168.1.11, network interfaceeth0 - Node 02 (Backup): IP
192.168.1.12, network interfaceeth0 - Virtual IP (Shared VIP): IP
192.168.1.100(the IP address exposed to clients or targeted by DNS)
Step 1: Install Keepalived
Install the required packages directly from the APT repository on both nodes:
sudo apt update
sudo apt install -y keepalived psmisc
Note: The psmisc package provides the killall utility, required by the Nginx process check script in the next step.
Step 2: Enable Non-local IP Binding
By default, Linux prevents processes from binding to an IP address that is not yet assigned to a local network interface. If Node 02 does not currently hold the VIP, Nginx will fail to start on that machine. Enable the ip_nonlocal_bind sysctl flag to resolve this:
echo "net.ipv4.ip_nonlocal_bind=1" | sudo tee -a /etc/sysctl.conf
sudo sysctl -p
Step 3: Configure the Master Node (192.168.1.11)
Create a new primary configuration file at /etc/keepalived/keepalived.conf:
sudo nano /etc/keepalived/keepalived.conf
Configuration content:
global_defs {
router_id lb01
enable_script_security
script_user root
}
vrrp_script check_nginx {
script "/usr/bin/killall -0 nginx"
interval 2
weight -20
fall 2
rise 2
}
vrrp_instance VI_WEB {
state MASTER
interface eth0
virtual_router_id 51
priority 101
advert_int 1
authentication {
auth_type PASS
auth_pass M@tKhauB@oMat123
}
virtual_ipaddress {
192.168.1.100/24 dev eth0 label eth0:vip
}
track_script {
check_nginx
}
}
Key points to note:
virtual_router_id 51: The unique identifier for the VRRP group. This value must match on both Master and Backup (range: 1 to 255).priority 101: Node priority level. The node with the higher priority holds the VIP.vrrp_script check_nginx: Checks the Nginx process every 2 seconds. If Nginx stops, priority drops by 20 points (down to 81). Since 81 is lower than the Backup’s 100, a failover is triggered immediately.
Step 4: Configure the Backup Node (192.168.1.12)
Create the same configuration file on Node 02:
sudo nano /etc/keepalived/keepalived.conf
File content:
global_defs {
router_id lb02
enable_script_security
script_user root
}
vrrp_script check_nginx {
script "/usr/bin/killall -0 nginx"
interval 2
weight -20
fall 2
rise 2
}
vrrp_instance VI_WEB {
state BACKUP
interface eth0
virtual_router_id 51
priority 100
advert_int 1
authentication {
auth_type PASS
auth_pass M@tKhauB@oMat123
}
virtual_ipaddress {
192.168.1.100/24 dev eth0 label eth0:vip
}
track_script {
check_nginx
}
}
The only differences: state changes to BACKUP and priority drops to 100.
Step 5: Start the Service and Verify Status
Enable and start the Keepalived service on both machines:
sudo systemctl enable --now keepalived
Check the network interface on Node 01 (Master):
ip addr show eth0
The secondary IP line 192.168.1.100/24 will appear with the label eth0:vip. Running the same command on the Backup node will show only its original static IP.
Testing the Failover Scenario
Open a terminal on a client machine and send continuous pings to the VIP:
ping 192.168.1.100
On Node 01, manually stop Nginx to trigger the failure condition:
sudo systemctl stop nginx
Observe the ping terminal. You will notice only a single dropped packet (roughly 1 second of timeout). Verify with ip addr show eth0 on Node 02: the VIP 192.168.1.100 has been assigned to the backup machine. Client HTTP requests continue to receive 200 OK responses as usual.
Once you restart Nginx on Node 01 (sudo systemctl start nginx), the Master’s priority returns to 101. The VIP is immediately reclaimed by Node 01 via preemption. In production environments, pairing your proxy tier with proactive monitoring like Prometheus with Grafana on Ubuntu Server ensures you track these failover events in real time. Your system is now resilient and self-healing against gateway hardware failures.

