Configuring LUKS and Clevis/Tang on CentOS Stream 9: Automatic Disk Encryption (NBDE)

CentOS tutorial - IT technology blog
CentOS tutorial - IT technology blog

After the CentOS 8 EOL situation, I had to hastily migrate 5 servers to Rocky Linux within a week. One of my biggest regrets was not having NBDE set up before the migration — which meant rebuilding everything from scratch on the new environment. Learning from that experience, this time I’m documenting the full process for CentOS Stream 9.

Background: The Real-World Problem with Traditional LUKS

Imagine you have a bare-metal server in a datacenter with LUKS-encrypted disks. Every time the server restarts — whether for a kernel update, power outage, or maintenance — you have to sit at the console and manually enter the passphrase. With 5–10 servers that’s already exhausting, let alone a larger fleet or an environment without overnight staff.

NBDE (Network-Bound Disk Encryption) solves exactly this problem: disks unlock automatically when the server boots within the internal network (where a Tang server is present), but still require a manual passphrase if the server is removed from the network. That’s the elegant part — physical security without blocking automation.

NBDE is best suited for:

  • Database servers storing sensitive data (PII, financial data)
  • Bare-metal or VMs at co-location datacenters
  • Compliance requirements: HIPAA, PCI-DSS, ISO 27001
  • Servers that auto-reboot after unattended upgrades without requiring staff on duty

The architecture consists of two main components:

  • Tang server: A dedicated key server using the McTLS protocol. Tang does not store the client’s key — it only participates in the key derivation process via Diffie-Hellman. Even if the server is compromised, the LUKS key is not exposed.
  • Clevis: A client-side framework supporting multiple "pins" (tang, tpm2, sss). Binds a LUKS volume to Tang or TPM.

Boot flow: Dracut initramfs → Clevis connects to Tang over the network → Tang returns data for key derivation → LUKS unlocks → boot continues.

Installing Tang and Clevis

1. Installing the Tang Server

Tang should run on a separate server or VM — not on the server you need to unlock. On CentOS Stream 9:

dnf install tang -y
systemctl enable --now tangd.socket
firewall-cmd --add-service=tangd --permanent
firewall-cmd --reload

Check that Tang is listening and retrieve the thumbprint (needed when binding Clevis):

systemctl status tangd.socket

# Get the key thumbprint
jose jwk thp -i /var/db/tang/*.jwk

Save the thumbprint — for example: abc123XYZdef456. You’ll need it in the binding step.

2. Installing Clevis on the Client

On the CentOS Stream 9 server that needs automatic encryption:

dnf install clevis clevis-luks clevis-dracut -y

# Optional: add TPM2 support
dnf install clevis-tpm2 -y

Detailed Configuration

Step 1: Identify the LUKS Volume to Bind

# Check LUKS devices
lsblk -f | grep crypto
blkid | grep LUKS

# View LUKS header info (slots in use, UUID, etc.)
cryptsetup luksDump /dev/sda3

I prefer using UUIDs instead of device names to avoid confusion when adding new disks. Save the UUID from the blkid output.

Step 2: Bind LUKS to Tang

# Replace TANG_SERVER_IP and THUMBPRINT with your actual values
clevis luks bind -d /dev/sda3 tang \
  '{"url":"http://TANG_SERVER_IP","thp":"THUMBPRINT"}'

Clevis will prompt for the current LUKS passphrase to add a new slot. After success, LUKS will have 2 slots: slot 0 (original passphrase) and slot 1 (Clevis/Tang). Verify:

clevis luks list -d /dev/sda3
# Output:
# 1: tang '{"url":"http://192.168.1.50","thp":"abc123..."}'

Step 3: Rebuild initramfs (the Most Often Skipped Step)

This is the step I once skipped, and the result was sitting at the passphrase prompt at 2 AM after rebooting a production server. You must rebuild initramfs to include Clevis in the bootloader:

dracut -fv --regenerate-all

Confirm Clevis is included in the initramfs:

lsinitrd /boot/initramfs-$(uname -r).img | grep clevis

Step 4: Update /etc/crypttab

For a data volume (not root), add to /etc/crypttab:

# Format: name  device  keyfile  options
data-encrypted  UUID=your-luks-uuid  none  _netdev,x-initrd.attach

The _netdev option tells systemd that networking is required before mounting. x-initrd.attach ensures it’s processed within the initrd. Without both options, the volume won’t unlock in the correct boot order.

Step 5 (Optional): High Availability with SSS — 2 Tang Servers

For high availability, SSS (Shamir’s Secret Sharing) allows combining multiple pins with a threshold. For example, requiring at least 1 out of 2 Tang servers:

clevis luks bind -d /dev/sda3 sss \
  '{"t":1,"pins":{"tang":[{"url":"http://tang1.internal","thp":"THUMB1"},{"url":"http://tang2.internal","thp":"THUMB2"}]}}'

This setup allows the server to boot even when one of the two Tang servers is down.

Testing and Monitoring

Testing Without a Reboot

# Test if Clevis can decrypt
clevis luks unlock -d /dev/sda3 -n test-unlock
# On success → creates /dev/mapper/test-unlock

# Clean up after testing
cryptsetup close test-unlock

# Test direct Tang connection
curl -sS http://TANG_SERVER_IP/adv | jose fmt -j- -g payload -y -o-
# Output is JSON containing keys — if empty, check the firewall

Viewing Boot Logs

# Clevis logs from the current boot
journalctl -b 0 | grep -i clevis

# Tang server logs in real-time
journalctl -u tangd -f

Monitoring Tang with Prometheus Blackbox

Tang has no built-in metrics. I use Blackbox Exporter to check the /adv endpoint:

# blackbox.yml
modules:
  tang_check:
    prober: http
    http:
      valid_status_codes: [200]
      method: GET
      fail_if_body_not_matches_regexp:
        - "keys"
# prometheus.yml
- job_name: 'tang'
  metrics_path: /probe
  params:
    module: [tang_check]
  static_configs:
    - targets:
      - http://TANG_SERVER_IP/adv
  relabel_configs:
    - source_labels: [__address__]
      target_label: __param_target
    - target_label: __address__
      replacement: localhost:9115

Critical alert: if Tang is down for more than 5 minutes, page on-call immediately — because the next reboot will require a manual passphrase.

Rotating the Tang Key Periodically

# On the Tang server: generate a new key
tangd-keygen /var/db/tang

# On the client: re-bind with the new key (slot 1)
clevis luks regen -d /dev/sda3 -s 1

I schedule key rotation every 6 months, triggering re-binding via an Ansible playbook so I don’t have to SSH into each server manually.

Backing Up the LUKS Header — Mandatory

After the CentOS 8 to Rocky Linux migration I dealt with earlier, I learned this the hard way: back up the LUKS header immediately after configuration. A lost header means permanently lost data — there is no recovery:

cryptsetup luksHeaderBackup /dev/sda3 \
  --header-backup-file /root/luks-header-sda3-backup.img

# Store this file somewhere else (not on the same server)
scp /root/luks-header-sda3-backup.img backup-server:/secure-backups/

Handling dracut Timeout Errors

If boot times out because Clevis can’t reach Tang in time, increase the timeout in the dracut config:

# /etc/dracut.conf.d/clevis.conf
kernel_cmdline="rd.timeout=90 rd.neednet=1"

# Rebuild initramfs
dracut -fv --regenerate-all

Setting up NBDE takes about 1–2 hours the first time, but after that your entire fleet can restart automatically without anyone on duty — and the security team can rest easy knowing data-at-rest is properly encrypted.

Share: