Configuring SR-IOV on Linux: Achieving Near Bare-Metal Network Performance for KVM Virtual Machines

Network tutorial - IT technology blog
Network tutorial - IT technology blog

Comparing Network Provisioning Solutions for KVM Virtual Machines

Hypervisor-level network bottlenecks are the biggest hurdle when running high-load databases or network gateways processing millions of packets per second on KVM. Currently, the Linux ecosystem provides three primary network provisioning mechanisms for virtual machines:

  • Linux Bridge / Open vSwitch (VirtIO-Net): The default virtualized networking model. Packets must traverse the host OS network stack and vhost-net before reaching the VM. While flexible and seamlessly supporting Live Migration, it introduces 25–50 µs of latency due to continuous CPU context switching on the host.
  • Full PCIe NIC Passthrough: Directly hands over an entire physical network card to a single VM. It delivers 100% bare-metal performance. However, this approach is cost-inefficient because each physical port can only serve a single virtual machine.
  • SR-IOV (Single Root I/O Virtualization): A PCI-SIG standard extension that partitions a physical port (Physical Function – PF) into dozens of independent virtual cards at the hardware level (Virtual Functions – VF). Each VF is assigned directly to a VM via IOMMU/VFIO. Packets flow straight to the hardware, completely bypassing the host kernel.

Detailed Technical Specification Comparison

The table below summarizes real-world operational criteria across the three options:

Criteria VirtIO-Net (Bridge) PCIe Passthrough SR-IOV (VFIO)
Throughput & PPS Bottlenecked by host CPU clock speed 100% Line-rate Reaches ~98-99% hardware line rate
Latency High (25 – 50 µs) Ultra-low (< 2 µs) Very low (1.5 – 3 µs)
Host CPU Overhead Consumes 15-30% at 5-10 Gbps traffic Near 0% Near 0%
Shared VM Density Unlimited 1 VM / 1 Physical Port 8 – 64+ VFs depending on NIC chipset
Live Migration Supported natively in KVM Not supported Requires failover bonding with VirtIO

When Does Your Project Truly Need SR-IOV?

For standard web servers or REST APIs, VirtIO-Net remains the most balanced and convenient choice. You should only transition to SR-IOV in the following scenarios:

  • High-packet-density network applications: 5G Core (UPF) clusters, WebRTC media gateways, public DNS resolvers, or financial trading platforms requiring over 10 million PPS.
  • Distributed database systems: Redis Cluster, ScyllaDB, or PostgreSQL setups with continuous replication demanding predictable, jitter-free microsecond latency.
  • Infrastructure cost optimization: Utilizing a single Intel X710/E810 25GbE card to provide high-speed networking for 16–32 VMs simultaneously without bottlenecking host CPU resources.

Step-by-Step SR-IOV Configuration on Linux KVM

Step 1: Enable IOMMU in BIOS and Linux Kernel

Access your server’s BIOS/UEFI settings. Enable Intel VT-d (or AMD IOMMU) along with the SR-IOV Global Enable flag.

Next, configure kernel parameters on the Host OS by editing GRUB:

sudo nano /etc/default/grub

Add the IOMMU parameters to the GRUB_CMDLINE_LINUX_DEFAULT variable. The iommu=pt flag improves performance by bypassing address translation for non-passthrough devices:

# For Intel CPUs
GRUB_CMDLINE_LINUX_DEFAULT="quiet splash intel_iommu=on iommu=pt"

# For AMD CPUs
# GRUB_CMDLINE_LINUX_DEFAULT="quiet splash amd_iommu=on iommu=pt"

Update GRUB and reboot the server:

sudo update-grub
sudo reboot

Once the host is back online, verify that IOMMU is active and functioning properly:

dmesg | grep -E "DMAR|IOMMU"

Step 2: Allocate Virtual Functions (VFs) from the Physical Network Port

Identify the physical network interface you want to enable SR-IOV on (e.g., enp4s0f0):

ip link show enp4s0f0

Check the maximum number of VFs supported by the NIC controller:

cat /sys/class/net/enp4s0f0/device/sriov_totalvfs

Initialize 4 Virtual Functions on this port:

echo 4 | sudo tee /sys/class/net/enp4s0f0/device/sriov_numvfs

Verify the newly created virtual interfaces at the PCIe layer:

lspci | grep -i ethernet

Step 3: Assign Static MAC Addresses and Enable Trust Mode

Assigning a static MAC address directly on the host prevents VFs from receiving random MAC addresses upon reboot. Additionally, setting trust on allows the VM to run in promiscuous mode or seamlessly integrate DPDK:

# Assign static MAC and enable Trust mode for VF 0
sudo ip link set enp4s0f0 vf 0 mac 52:54:00:11:22:33 trust on

# (Optional) Assign VLAN tag 100 if the infrastructure uses trunking
sudo ip link set enp4s0f0 vf 0 vlan 100

Step 4: Attach the VF to the KVM Virtual Machine via Libvirt

Locate the PCI bus address of the configured VF using lspci (e.g., 0000:04:10.0):

lspci -s 04:10.0

Open the XML configuration file of the target virtual machine (e.g., db-prod-vm):

virsh edit db-prod-vm

Add the <interface type='hostdev'> block inside the <devices> section:

<interface type='hostdev' managed='yes'>
  <source>
    <address type='pci' domain='0x0000' bus='0x04' slot='0x10' function='0x0'/>
  </source>
  <mac address='52:54:00:11:22:33'/>
</interface>

Start the virtual machine and open the console to verify:

virsh start db-prod-vm
virsh console db-prod-vm

Inside the guest OS, running ip a will show the new network interface powered by native vendor drivers (such as iavf for Intel 700/800 series or ixgbevf for X520/X540 series).

3 Critical Lessons Learned Running SR-IOV in Production

In production deployments, SR-IOV setups can occasionally encounter intermittent packet loss during peak hours even when bandwidth limits have not been reached. Here is how to resolve the 3 most common issues:

  1. Increase Ring Buffer Sizes in the Guest VM: By default, VF drivers only allocate 256 to 512 descriptors for RX/TX buffers. During micro-burst traffic spikes, the NIC drops packets directly in hardware queues. Maximize the RX/TX ring buffers inside the VM (typically 4096):
# View current ring buffer limits
sudo ethtool -g eth1

# Max out the ring buffer sizes
sudo ethtool -G eth1 rx 4096 tx 4096
  1. Disable Spoof Checking for VIP / Keepalived: If your VM runs a Virtual IP (Keepalived, VRRP, CARP), the physical NIC will drop packets because anti-spoofing security filters detect unfamiliar MAC/IP combinations. Disable spoof checking on the host for that specific VF:
sudo ip link set enp4s0f0 vf 0 spoofchk off
  1. Persist VF Initialization Across Reboots: The sriov_numvfs value resets to 0 whenever the host reboots. The most reliable solution is creating a oneshot systemd service to automatically re-instantiate the VFs:
# Create service file /etc/systemd/system/sriov.service
[Unit]
Description=Automatically initialize SR-IOV VFs at boot
After=network.target

[Service]
Type=oneshot
ExecStart=/bin/sh -c 'echo 4 > /sys/class/net/enp4s0f0/device/sriov_numvfs'
RemainAfterExit=yes

[Install]
WantedBy=multi-user.target

Enable and start the service to make the configuration permanent:

sudo systemctl enable --now sriov.service
Share: