Zero-Downtime KVM/Proxmox Live Migration: Dedicated Networks, ZSTD Compression, and Post-Copy

Virtualization tutorial - IT technology blog
Virtualization tutorial - IT technology blog

1. The VM Migration Nightmare: Stuck at 95% and Saturated Cluster Networks

Many sysadmins are all too familiar with this scenario: A physical server hosting a production PostgreSQL cluster or an e-commerce web app needs a faulty RAM stick replaced. Confident in your setup, you trigger a Live Migration from Node A to Node B, expecting zero service interruption.

Reality, however, hits hard. The migration crawls up to 95% and stalls indefinitely. The culprit? The guest workload writes to RAM faster than the network adapter can transfer dirty memory pages to the destination host (known as migration non-convergence).

To make matters worse, the flood of RAM data completely saturates the shared network interface. Internal packets start dropping rapidly, staff lose connection to the CRM, and external users are immediately greeted with 504 Gateway Timeout errors.

Even in my small lab cluster with 12 Proxmox VMs, I encountered this exact issue. Whenever I migrated a write-heavy Redis or PostgreSQL VM generating 150–200 MB/s of dirty memory over a 1Gbps link, the cluster network choked completely. Ensuring high availability and stability in production requires addressing the root causes directly.

2. Technical Deep Dive: Why Does Live Migration Fail?

To put it simply: Live Migration is like trying to move your belongings to a new house while someone back at the old house keeps buying new furniture faster than you can haul it away.

The Pre-Copy Loop and Dirty Memory

By default, KVM/QEMU and Proxmox VE employ a Pre-copy Migration mechanism structured in three stages:

  1. Stage 1 (Initial Copy): Transfer the VM’s entire memory footprint from Node A to Node B while the VM continues serving live requests.
  2. Stage 2 (Iterative Phase): During the transfer, the VM continuously modifies memory pages (known as Dirty Pages). KVM gathers these dirty pages and sends them over in successive rounds.
  3. Stage 3 (Cutover): Once the remaining volume of dirty pages drops below a minimal threshold, KVM pauses the VM on Node A for about 50–100ms, pushes the final few megabytes, and resumes execution on Node B.

Failure occurs when the Memory Dirty Rate > Network Bandwidth. KVM gets trapped in an infinite iterative loop: bandwidth gets exhausted, network switches get overwhelmed, and the migration never converges.

Network Contention

Routing migration traffic over the same network interface as application traffic and Corosync (cluster heartbeat) introduces even greater risks. When migration saturates 100% of the physical link, Corosync heartbeats time out. The cluster assumes the node has dropped offline and triggers fencing mechanisms (force-rebooting the host). What started as routine maintenance quickly spirals into a full-scale cluster outage.

3. Common Workarounds (and Their Limitations)

Option 1: Rate Limiting Migration Traffic

The quickest tactical fix is throttling migration bandwidth to protect the shared network link:

# Limit migration speed to a maximum of 50 MB/s on KVM
virsh migrate-setspeed <vm_name> 50

Drawback: While this shields the internal network, it significantly lengthens migration duration. For VMs with high write rates, bandwidth throttling practically guarantees the migration will never finish.

Option 2: Throttling Guest vCPU (Auto-Converge)

When the hypervisor detects that dirty pages aren’t decreasing across iterations, it automatically throttles the VM’s vCPU cycles to slow down its memory dirtying rate.

# Enable auto-converge on KVM via virsh
virsh migrate --live --auto-converge <vm_name> qemu+ssh://10.10.10.2/system

Drawback: vCPU performance can be throttled anywhere from 20% up to 80%. Hosted applications (especially databases) will experience severe latency spikes, directly impacting end users.

4. The Production-Grade Strategy: Dedicated Networks, Data Compression, and Post-Copy

To achieve truly seamless, zero-downtime migrations without service degradation, implement the following robust architecture:

Step 1: Isolate a Dedicated Migration Network (10Gbps+ Recommended)

Dedicate a physical NIC (or a bonded pair of independent interfaces) exclusively for inter-node migration traffic.

Suppose you have two nodes configured on a dedicated subnet 10.10.10.0/24 via eth1:

  • Node A: 10.10.10.1/24
  • Node B: 10.10.10.2/24

On Proxmox VE, configure this directly via the GUI or edit /etc/pve/datacenter.cfg:

# Define the dedicated migration network in datacenter.cfg
migration: type=secure,network=10.10.10.0/24

On standalone KVM (libvirt), explicitly specify the migration endpoint IP:

# Execute live migration targeting the dedicated network IP
virsh migrate --live --verbose \
  --migrateuri tcp://10.10.10.2:49152 \
  my-production-vm qemu+ssh://10.10.10.2/system

Step 2: Enable Stream Compression (ZSTD)

Compressing memory pages before sending them across the wire reduces payload size by 30% to 60%. This is particularly effective for VMs with unallocated RAM or repetitive in-memory data.

In Proxmox VE, navigate to Datacenter -> Options -> Migration Settings and set the compression algorithm to ZSTD to achieve the optimal balance between CPU utilization and compression throughput.

Step 3: Use Post-Copy for Heavy Workloads

When dealing with write-intensive databases (like MySQL, PostgreSQL, or Redis) where Pre-copy struggles to converge, Post-copy offers the definitive solution.

Unlike Pre-copy, Post-copy operates under a different paradigm:

  1. Initial baseline memory is transferred using standard Pre-copy.
  2. If migration fails to converge, the hypervisor pauses the VM on Node A and immediately starts it on Node B (with cutover downtime lasting only tens of milliseconds).
  3. The VM resumes execution on Node B right away. When it accesses a memory page that has not yet migrated, Node B triggers a remote Userfaultfd (Page Fault) to pull the requested page on-demand from Node A.
  4. Node A concurrently streams the remaining background RAM pages until the entire state is transferred.

Executing Post-copy in KVM/libvirt:

# 1. Start live migration with post-copy capability enabled
virsh migrate --live --postcopy --verbose my-database-vm qemu+ssh://10.10.10.2/system

# 2. If the migration loop drags on, trigger the post-copy transition immediately:
virsh migrate-postcopy my-database-vm

Combining a dedicated 10Gbps network, ZSTD compression, and the Post-copy technique gives you the confidence to migrate any high-load virtual machine without risking network saturation or service outages.

Share: