The High Availability Trap: When Two Redundant VMs Land on the Same Physical Host
Your infrastructure runs two web virtual machines in parallel behind a load balancer. In theory, if one node fails, the other seamlessly shoulders the workload. Everything appears to meet High Availability (HA) standards.
Then disaster strikes. An ESXi host suffers a sudden power loss due to a failed power supply unit (PSU). The entire web service immediately goes down. Opening vCenter to investigate, you discover that both redundant virtual machines were running on that single failed physical host all along. The primary culprit is Distributed Resource Scheduler (DRS). Seeing abundant free RAM on that host, DRS consolidated both VMs onto it to optimize resource utilization.
Across an ESXi cluster of 4 to 16 hosts, DRS calculates placement based solely on CPU and RAM consumption on a 5-minute cycle. The system has no inherent awareness of which virtual machines serve as failover counterparts for one another. To eliminate this critical vulnerability, you must implement DRS Affinity Rules and Anti-Affinity Rules.
Understanding the Basics: What Are Affinity and Anti-Affinity Rules?
The core operating principle is straightforward:
- Affinity (Attract): Enforces or prioritizes grouping specified objects together on the same host.
- Anti-Affinity (Repel): Mandates separating specified objects across different hosts, strictly preventing them from coexisting on the same physical server.
1. VM-VM Rules (Inter-VM Placement Rules)
- Keep Virtual Machines Together (Affinity): Keeps a designated group of VMs running on the same ESXi host. A classic example is an application and database VM pair handling thousands of internal requests per second. Co-locating them allows traffic to route directly through the in-memory vSwitch, achieving multi-gigabit throughput without consuming physical top-of-rack switch bandwidth.
- Separate Virtual Machines (Anti-Affinity): Requires specified VMs to reside on distinct physical hosts. This is a vital rule for paired Active Directory Domain Controllers, SQL Server AlwaysOn availability groups, or redundant NGINX reverse proxies.
2. VM-Host Rules (Mapping VM Groups to Host Groups)
This mechanism binds a designated group of virtual machines (VM DRS Group) to a specific group of physical servers (Host DRS Group). It operates under two enforcement levels:
- Must run on / Must NOT run on (Hard Rule – Strict Enforcement): Strictly enforced by both DRS and vSphere HA. If all hosts in the target group fail, the VMs will remain powered off rather than restarting on unauthorized hosts. This level is primarily used for core-based software licensing compliance (such as Oracle DB or MS SQL Server Enterprise) or servers with dedicated physical hardware attachments (PCIe Passthrough, GPUs).
- Should run on / Should NOT run on (Soft Rule – Preferential Placement): DRS keeps VMs on the designated host group under normal conditions. However, during host outages or resource shortages, vSphere HA is permitted to power on the VMs on alternative hosts to preserve service uptime.
Step-by-Step Guide: Configuring Rules in vSphere
Method 1: Creating a VM-VM Anti-Affinity Rule via vSphere Client
Example scenario: Separating VM-Web-01 and VM-Web-02 across different ESXi hosts.
- Log in to the vSphere Client (HTML5 interface).
- Select the target Cluster in the inventory navigation tree.
- Navigate to the Configure tab > select VM/Host Rules.
- Click Add to create a new rule.
- Enter a rule name:
Rule-AntiAffinity-WebTier. - Check the Enable rule checkbox.
- In the Type dropdown, select
Separate Virtual Machines. - Click Add in the VM list, select
VM-Web-01andVM-Web-02. - Click OK. If both VMs currently reside on the same physical host, DRS will immediately trigger a vMotion migration to move one VM to an alternate host.
Method 2: Configuring a VM-Host Rule for Database VMs (License Optimization)
Example scenario: In an 8-host cluster, Oracle licenses are only purchased for 2 hosts (esxi-01 and esxi-02). We need to restrict 3 database VMs to run strictly on these 2 licensed nodes.
- Navigate to the Cluster’s Configure tab > select VM/Host Groups.
- Create a VM Group: Name it
VMG-Oracle-DBsand select the 3 database VMs. - Create a Host Group: Name it
HG-Oracle-Licensed-Hostsand selectesxi-01andesxi-02. - Switch to the VM/Host Rules section and click Add.
- In the Type dropdown, select
Virtual Machines to Hosts. - Set VM Group: to
VMG-Oracle-DBs. - Select the relationship:
Should run on hosts in group(recommended for high availability failover) orMust run on hosts in group(if strict license audit compliance is mandatory). - Set Host Group: to
HG-Oracle-Licensed-Hostsand click OK.
Method 3: Automation Using VMware PowerCLI
Use PowerShell scripting when managing multiple clusters or integrating into VM provisioning CI/CD pipelines:
# 1. Connect to vCenter Server
Connect-VIServer -Server vcenter.company.local -User "[email protected]" -Password "MatKhau@123"
# 2. Define parameters
$clusterName = "Production-Cluster"
$ruleName = "AntiAffinity-DomainControllers"
$vmList = Get-VM -Name "DC-01", "DC-02"
# 3. Create VM-VM Anti-Affinity Rule
$cluster = Get-Cluster -Name $clusterName
New-DrsRule -Cluster $cluster -Name $ruleName -KeepTogether $false -VM $vmList -Enabled $true
# 4. Verify active DRS rules
Get-DrsRule -Cluster $cluster | Format-Table Name, Enabled, KeepTogether
Script to create a VM-Host Rule (Soft Rule):
# Create VM Group and Host Group
$vmGroup = New-DrsClusterGroup -Cluster $cluster -Name "VMG-AppServers" -VM (Get-VM -Name "App-01", "App-02")
$hostGroup = New-DrsClusterGroup -Cluster $cluster -Name "HG-Rack01" -VMHost (Get-VMHost -Name "esxi-01.local", "esxi-02.local")
# Create Soft Rule (ShouldRunOn)
New-DrsRule -Cluster $cluster -Name "Rule-App-To-Rack01" -Type VmHostRule `
-VMGroup $vmGroup -HostGroup $hostGroup -RuleType "ShouldRunOn" -Enabled $true
3 Common Pitfalls That Can Paralyze Your Infrastructure When Using DRS Rules
- Maintenance Mode Stuck at 66%: Entering Maintenance Mode on an ESXi host triggers DRS to evacuate all resident VMs. If a host carries a VM bound by a Must run on rule and the remaining target hosts lack sufficient CPU/RAM, the maintenance task will hang indefinitely at 66%.
- Conflicting Rules: Setting Rule A to co-locate
VM-01withVM-02while Rule B prevents them from running on the same host group creates an unresolvable loop. When a conflict occurs, vCenter triggers a yellow alert and completely disables automatic DRS balancing for the impacted virtual machines. - Overusing “Must” Instead of “Should”: site driven by strict per-core licensing compliance or hardware constraints, always favor Should run on soft rules. Enforcing a hard Must rule strips vSphere HA of its ability to automatically recover your workloads during physical hardware outages.
Summary
Robust enterprise hardware alone does not guarantee infrastructure resilience. The key lies in intelligent VM placement strategy. Taking 5 minutes to audit and configure Anti-Affinity Rules for critical workloads will permanently protect your services from total outages caused by a single hardware component failure.

