The Real Problem: Single-Node RabbitMQ and Its Single Point of Failure
I once watched an e-commerce order processing system grind to a complete halt because the server running single-node RabbitMQ got OOM-killed at 2 AM. Every message pending in the queue vanished, 200+ consumer processes had to be restarted manually, and the dev team spent nearly three hours getting everything back to normal.
This wasn’t a RabbitMQ bug — it was an architectural problem. With a single point of failure, the question isn’t “will it go down” but “when will it go down.” RabbitMQ clustering solves this by replicating queue metadata across multiple nodes. With Quorum Queues, even message data is replicated — no messages are lost even if a node dies mid-flight.
I’ve been using Fedora as my primary development machine for two years and I genuinely appreciate how fast the packages get updated. But when deploying to production on Fedora Server, SELinux and firewalld have a habit of causing problems without any warning. RabbitMQ is particularly tricky because it needs to open several unusual ports and write to a variety of paths — things that the default SELinux policy doesn’t fully cover.
Three Common Ways to Deploy a RabbitMQ Cluster
Option 1: Bare-Metal Cluster (Install Directly on the OS)
Install the rabbitmq-server package directly on each server, sync the Erlang cookie, then join the cluster using rabbitmqctl join_cluster. This approach gives you full control over every detail — from file descriptor limits and memory watermarks to the SELinux context of individual files.
Pros: No container overhead, easy to monitor via systemd/journald, integrates well with Prometheus node_exporter. Erlang 26 is available in the official Fedora 40 repo — no need to add EPEL or third-party repos like you would on CentOS Stream.
Cons: You have to handle SELinux and firewalld manually. Rolling upgrades require more care. Harder to reproduce the dev environment.
Option 2: Podman Cluster with Persistent Volumes
Use the rabbitmq:3-management image, create a pod, or use Podman Compose. Podman is the default choice on Fedora — it handles rootless containers better than Docker and doesn’t require a background daemon.
Pros: Portable, easy to upgrade by pulling a new image, no need to deal with SELinux at the application level since containers provide isolation out of the box.
Cons: You still need to configure firewalld for exposed ports. Networking across multiple Podman hosts requires an extra layer — a VPN or overlay network. Persistent volumes with SELinux label :z are easy to forget and cause permission errors that are a nightmare to debug.
Option 3: Kubernetes with the RabbitMQ Cluster Operator
Deploy using the RabbitMQ Cluster Operator via Helm. The Operator automatically manages cluster joins and departures, handles rolling upgrades, and scales in and out.
Pros: Fully automated, self-healing, integrates seamlessly with the Kubernetes monitoring stack.
Cons: Kubernetes overhead is significant — you need at least three worker nodes just to run the control plane properly. For small-to-medium workloads on bare-metal VPS, this is clear overkill.
Which Option for Fedora Server?
It depends on your current infrastructure:
- 2–3 bare-metal VPS servers, want maximum control, no Kubernetes → Bare-metal cluster
- Already have Podman infrastructure, need portability and easy reproducibility → Podman cluster
- Running Kubernetes, want to integrate into the same ecosystem → K8s Operator
I’ll walk through the bare-metal cluster — the most common scenario for Fedora Server on VPS, and the one where SELinux and firewalld cause the most friction if you don’t configure them correctly from the start.
Deploying a 3-Node RabbitMQ Cluster on Fedora Server
Step 0: Set Up Hostnames and /etc/hosts
You’ll need three Fedora Server 40+ machines. I’m using:
rmq1— 192.168.1.10rmq2— 192.168.1.11rmq3— 192.168.1.12
Add the following to /etc/hosts on all three nodes. RabbitMQ cluster identifies nodes by hostname, not IP — getting this wrong is the root cause of most “node not found” errors when joining a cluster:
sudo tee -a /etc/hosts <<EOF
192.168.1.10 rmq1
192.168.1.11 rmq2
192.168.1.12 rmq3
EOF
Set the hostname to match what you declared above:
# Run on each node accordingly
sudo hostnamectl set-hostname rmq1 # or rmq2, rmq3
Step 1: Install Erlang and RabbitMQ
Fedora 40 ships Erlang 26 in its official repo — no need to add EPEL or third-party repos like on CentOS Stream:
# Run on all 3 nodes
sudo dnf install -y erlang rabbitmq-server
# Check versions
erl -version
rabbitmqctl version
# Enable and start the service
sudo systemctl enable --now rabbitmq-server
Step 2: Sync the Erlang Cookie — the Most Commonly Skipped Step
Erlang nodes can only communicate when they share the same Erlang cookie — a string stored at /var/lib/rabbitmq/.erlang.cookie. A cookie mismatch produces an “authentication failed” error that’s very easy to confuse with a network or firewall issue — this is the most commonly skipped step when setting up a cluster for the first time.
# On rmq1: read the current cookie
sudo cat /var/lib/rabbitmq/.erlang.cookie
# Example output: WKDXQABCMNOPQRSTUVWX
# Stop the service before making changes
sudo systemctl stop rabbitmq-server
# On rmq2 and rmq3: overwrite with the cookie from rmq1
echo -n "WKDXQABCMNOPQRSTUVWX" | sudo tee /var/lib/rabbitmq/.erlang.cookie
sudo chmod 400 /var/lib/rabbitmq/.erlang.cookie
sudo chown rabbitmq:rabbitmq /var/lib/rabbitmq/.erlang.cookie
# Restart on all 3 nodes
sudo systemctl start rabbitmq-server
Step 3: Configure SELinux for RabbitMQ
Fedora doesn’t ship SELinux policies that cover all the ports a RabbitMQ cluster needs. Port 25672 (Erlang distribution) and 4369 (EPMD) are commonly blocked, and the audit log isn’t always clear about it — making it easy to mistake for a firewall problem.
First, check what SELinux is currently blocking:
sudo ausearch -m avc -ts recent | grep rabbitmq
# or
sudo journalctl -u rabbitmq-server | grep -i denied
Add the required ports to the amqp_port_t type:
# Install the SELinux port management tool
sudo dnf install -y policycoreutils-python-utils
# Add ports for AMQP, Management UI, Erlang distribution, and EPMD
sudo semanage port -a -t amqp_port_t -p tcp 5672
sudo semanage port -a -t amqp_port_t -p tcp 15672
sudo semanage port -a -t amqp_port_t -p tcp 25672
sudo semanage port -a -t amqp_port_t -p tcp 4369
# Verify
sudo semanage port -l | grep amqp
Still hitting permission errors after opening the ports? Generate a custom SELinux module from the audit log:
sudo ausearch -c 'beam.smp' --raw | audit2allow -M rabbitmq_custom
sudo semodule -i rabbitmq_custom.pp
# Verify the module is loaded
sudo semodule -l | grep rabbitmq
Step 4: Configure firewalld
A RabbitMQ cluster needs four main ports open between nodes:
- 4369/tcp — EPMD (Erlang Port Mapper Daemon) — node discovery
- 5672/tcp — AMQP protocol — consumer/producer connections
- 25672/tcp — Erlang distribution — inter-node cluster communication
- 15672/tcp — Management Web UI (restrict to trusted source IPs)
# Open cluster communication ports on all 3 nodes
sudo firewall-cmd --permanent --add-port=4369/tcp
sudo firewall-cmd --permanent --add-port=5672/tcp
sudo firewall-cmd --permanent --add-port=25672/tcp
# Management UI — restrict to internal network only
sudo firewall-cmd --permanent --add-rich-rule='rule family="ipv4" source address="192.168.1.0/24" port protocol="tcp" port="15672" accept'
# Reload
sudo firewall-cmd --reload
sudo firewall-cmd --list-ports
Step 5: Join the Cluster
With the Erlang cookie synced and ports open, join node2 and node3 into node1:
# On rmq2:
sudo rabbitmqctl stop_app
sudo rabbitmqctl reset
sudo rabbitmqctl join_cluster rabbit@rmq1
sudo rabbitmqctl start_app
# On rmq3:
sudo rabbitmqctl stop_app
sudo rabbitmqctl reset
sudo rabbitmqctl join_cluster rabbit@rmq1
sudo rabbitmqctl start_app
# Check cluster status from any node
sudo rabbitmqctl cluster_status
A successful output will list all three nodes under running_nodes: rabbit@rmq1, rabbit@rmq2, rabbit@rmq3.
Step 6: Create a Quorum Queue and Admin User
Classic Mirrored Queues have been deprecated since RabbitMQ 3.9. Use Quorum Queues instead — they use Raft consensus-based replication, which is far safer and more predictable. With a 3-node cluster, Raft requires a majority (2/3) to commit, so the queue keeps running even if one node goes down.
# Enable the Management Plugin
sudo rabbitmq-plugins enable rabbitmq_management
# Create an admin user — never use guest in production
sudo rabbitmqctl add_user admin StrongPassword123
sudo rabbitmqctl set_user_tags admin administrator
sudo rabbitmqctl set_permissions -p / admin ".*" ".*" ".*"
# Delete the default guest user
sudo rabbitmqctl delete_user guest
# Create a Quorum Queue via CLI
sudo rabbitmqadmin declare queue name=order_queue durable=true \
arguments='{"x-queue-type":"quorum"}'
Testing the Cluster: Publish from One Node, Consume from Another
Quick test with Python — publish to rmq2, consume from rmq3:
import pika
# Publish from rmq2
conn = pika.BlockingConnection(
pika.ConnectionParameters(
host='192.168.1.11',
credentials=pika.PlainCredentials('admin', 'StrongPassword123')
)
)
channel = conn.channel()
channel.queue_declare(
queue='order_queue', durable=True,
arguments={'x-queue-type': 'quorum'}
)
channel.basic_publish(exchange='', routing_key='order_queue', body='test-message')
conn.close()
print("Published")
Then consume from 192.168.1.12 (rmq3) with similar code. The message is still readable because Quorum Queues replicate data across all three nodes. Try shutting down rmq2 and consuming again — the queue keeps running without a hitch. That’s real HA, not just HA on paper.
Common Errors and How to Fix Them
“Node ‘rabbit@rmq2’ not running” when joining the cluster → Check the hostnames in /etc/hosts. A wrong hostname is the number one cause. Run sudo rabbitmqctl status to see what name the node is identifying itself as.
“Authentication failed” when joining → The Erlang cookie doesn’t match. Stop the service, copy the cookie from rmq1 again, set permissions to 400, and restart.
“Permission denied” even though firewalld ports are open → SELinux is blocking it, not the firewall. I once spent an entire afternoon debugging this because I assumed firewalld was the culprit, but it was actually SELinux preventing beam.smp from binding to port 25672. Lesson learned: always check /var/log/audit/audit.log directly instead of relying solely on ausearch — when the log gets rotated, ausearch -ts recent will miss the context you need.
Cluster split-brain after a network partition → Quorum Queues use Raft: the partition holding the majority of nodes (≥2/3) continues accepting writes, while the minority partition blocks itself. There’s no data inconsistency — this is the main reason to migrate from Classic Mirrored Queues to Quorum Queues.

