After the CentOS 8 EOL situation, I had to hastily migrate 5 servers to Rocky Linux within a week. One of my biggest regrets was not having NBDE set up before the migration — which meant rebuilding everything from scratch on the new environment. Learning from that experience, this time I’m documenting the full process for CentOS Stream 9.
Background: The Real-World Problem with Traditional LUKS
Imagine you have a bare-metal server in a datacenter with LUKS-encrypted disks. Every time the server restarts — whether for a kernel update, power outage, or maintenance — you have to sit at the console and manually enter the passphrase. With 5–10 servers that’s already exhausting, let alone a larger fleet or an environment without overnight staff.
NBDE (Network-Bound Disk Encryption) solves exactly this problem: disks unlock automatically when the server boots within the internal network (where a Tang server is present), but still require a manual passphrase if the server is removed from the network. That’s the elegant part — physical security without blocking automation.
NBDE is best suited for:
- Database servers storing sensitive data (PII, financial data)
- Bare-metal or VMs at co-location datacenters
- Compliance requirements: HIPAA, PCI-DSS, ISO 27001
- Servers that auto-reboot after unattended upgrades without requiring staff on duty
The architecture consists of two main components:
- Tang server: A dedicated key server using the McTLS protocol. Tang does not store the client’s key — it only participates in the key derivation process via Diffie-Hellman. Even if the server is compromised, the LUKS key is not exposed.
- Clevis: A client-side framework supporting multiple "pins" (tang, tpm2, sss). Binds a LUKS volume to Tang or TPM.
Boot flow: Dracut initramfs → Clevis connects to Tang over the network → Tang returns data for key derivation → LUKS unlocks → boot continues.
Installing Tang and Clevis
1. Installing the Tang Server
Tang should run on a separate server or VM — not on the server you need to unlock. On CentOS Stream 9:
dnf install tang -y
systemctl enable --now tangd.socket
firewall-cmd --add-service=tangd --permanent
firewall-cmd --reload
Check that Tang is listening and retrieve the thumbprint (needed when binding Clevis):
systemctl status tangd.socket
# Get the key thumbprint
jose jwk thp -i /var/db/tang/*.jwk
Save the thumbprint — for example: abc123XYZdef456. You’ll need it in the binding step.
2. Installing Clevis on the Client
On the CentOS Stream 9 server that needs automatic encryption:
dnf install clevis clevis-luks clevis-dracut -y
# Optional: add TPM2 support
dnf install clevis-tpm2 -y
Detailed Configuration
Step 1: Identify the LUKS Volume to Bind
# Check LUKS devices
lsblk -f | grep crypto
blkid | grep LUKS
# View LUKS header info (slots in use, UUID, etc.)
cryptsetup luksDump /dev/sda3
I prefer using UUIDs instead of device names to avoid confusion when adding new disks. Save the UUID from the blkid output.
Step 2: Bind LUKS to Tang
# Replace TANG_SERVER_IP and THUMBPRINT with your actual values
clevis luks bind -d /dev/sda3 tang \
'{"url":"http://TANG_SERVER_IP","thp":"THUMBPRINT"}'
Clevis will prompt for the current LUKS passphrase to add a new slot. After success, LUKS will have 2 slots: slot 0 (original passphrase) and slot 1 (Clevis/Tang). Verify:
clevis luks list -d /dev/sda3
# Output:
# 1: tang '{"url":"http://192.168.1.50","thp":"abc123..."}'
Step 3: Rebuild initramfs (the Most Often Skipped Step)
This is the step I once skipped, and the result was sitting at the passphrase prompt at 2 AM after rebooting a production server. You must rebuild initramfs to include Clevis in the bootloader:
dracut -fv --regenerate-all
Confirm Clevis is included in the initramfs:
lsinitrd /boot/initramfs-$(uname -r).img | grep clevis
Step 4: Update /etc/crypttab
For a data volume (not root), add to /etc/crypttab:
# Format: name device keyfile options
data-encrypted UUID=your-luks-uuid none _netdev,x-initrd.attach
The _netdev option tells systemd that networking is required before mounting. x-initrd.attach ensures it’s processed within the initrd. Without both options, the volume won’t unlock in the correct boot order.
Step 5 (Optional): High Availability with SSS — 2 Tang Servers
For high availability, SSS (Shamir’s Secret Sharing) allows combining multiple pins with a threshold. For example, requiring at least 1 out of 2 Tang servers:
clevis luks bind -d /dev/sda3 sss \
'{"t":1,"pins":{"tang":[{"url":"http://tang1.internal","thp":"THUMB1"},{"url":"http://tang2.internal","thp":"THUMB2"}]}}'
This setup allows the server to boot even when one of the two Tang servers is down.
Testing and Monitoring
Testing Without a Reboot
# Test if Clevis can decrypt
clevis luks unlock -d /dev/sda3 -n test-unlock
# On success → creates /dev/mapper/test-unlock
# Clean up after testing
cryptsetup close test-unlock
# Test direct Tang connection
curl -sS http://TANG_SERVER_IP/adv | jose fmt -j- -g payload -y -o-
# Output is JSON containing keys — if empty, check the firewall
Viewing Boot Logs
# Clevis logs from the current boot
journalctl -b 0 | grep -i clevis
# Tang server logs in real-time
journalctl -u tangd -f
Monitoring Tang with Prometheus Blackbox
Tang has no built-in metrics. I use Blackbox Exporter to check the /adv endpoint:
# blackbox.yml
modules:
tang_check:
prober: http
http:
valid_status_codes: [200]
method: GET
fail_if_body_not_matches_regexp:
- "keys"
# prometheus.yml
- job_name: 'tang'
metrics_path: /probe
params:
module: [tang_check]
static_configs:
- targets:
- http://TANG_SERVER_IP/adv
relabel_configs:
- source_labels: [__address__]
target_label: __param_target
- target_label: __address__
replacement: localhost:9115
Critical alert: if Tang is down for more than 5 minutes, page on-call immediately — because the next reboot will require a manual passphrase.
Rotating the Tang Key Periodically
# On the Tang server: generate a new key
tangd-keygen /var/db/tang
# On the client: re-bind with the new key (slot 1)
clevis luks regen -d /dev/sda3 -s 1
I schedule key rotation every 6 months, triggering re-binding via an Ansible playbook so I don’t have to SSH into each server manually.
Backing Up the LUKS Header — Mandatory
After the CentOS 8 to Rocky Linux migration I dealt with earlier, I learned this the hard way: back up the LUKS header immediately after configuration. A lost header means permanently lost data — there is no recovery:
cryptsetup luksHeaderBackup /dev/sda3 \
--header-backup-file /root/luks-header-sda3-backup.img
# Store this file somewhere else (not on the same server)
scp /root/luks-header-sda3-backup.img backup-server:/secure-backups/
Handling dracut Timeout Errors
If boot times out because Clevis can’t reach Tang in time, increase the timeout in the dracut config:
# /etc/dracut.conf.d/clevis.conf
kernel_cmdline="rd.timeout=90 rd.neednet=1"
# Rebuild initramfs
dracut -fv --regenerate-all
Setting up NBDE takes about 1–2 hours the first time, but after that your entire fleet can restart automatically without anyone on duty — and the security team can rest easy knowing data-at-rest is properly encrypted.

