A Guide to MySQL High Availability with Orchestrator: Automated Topology Discovery and Safe 2 AM Failover

MySQL tutorial - IT technology blog
MySQL tutorial - IT technology blog

At 2:00 AM, your PagerDuty alarm blares: CRITICAL - MySQL Master Unreachable. You scramble to open a terminal, attempt to SSH into the primary node, and hit a connection timeout triggered by a kernel panic. Meanwhile, the backend floods your monitoring with 500 errors as transactions grind to a halt.

Without automated High Availability (HA), you are facing 15 to 30 minutes of high-stress firefighting. You have to inspect every replica to see which one has the lowest replication lag, promote it, and manually update connection strings. With Orchestrator, this nerve-racking ordeal is resolved cleanly in under 15 seconds.

Three Common Approaches to MySQL High Availability

When implementing HA for a MySQL replication cluster, DevOps engineers typically evaluate three main solutions:

  • Keepalived / Corosync with Virtual IP (VIP): Sets up a dual-node master-master or active-passive cluster using a floating IP. Continuous heartbeats monitor node health and swing the IP over when an outage occurs.
  • MHA (Master High Availability): A Perl-based toolset that once served as the industry standard. MHA connects to each node via SSH, inspects state, applies missing relay logs to eliminate data loss, and promotes the best candidate replica to master.
  • Orchestrator: An open-source tool written in Go, originally developed by Shlomi Noach while at GitHub. It operates as an independent daemon communicating over standard MySQL port 3306, automatically mapping replication topologies and managing failovers based on replica consensus.

Detailed Comparison: Pros and Cons

1. Keepalived + Virtual IP

Pros: Very quick to set up, minimal resource footprint, and highly familiar to systems administrators, especially when configuring MySQL High Availability with Keepalived and HAProxy for basic failover needs.

Cons: Lacks any genuine database awareness. Keepalived only checks port 3306 availability or ICMP pings. It cannot detect whether a replica is lagging by 30 minutes or if the master is suffering from storage I/O bottlenecks. Its greatest danger is a split-brain scenario: both nodes accepting writes simultaneously, resulting in catastrophic data inconsistency and corruption.

2. MHA (Master High Availability)

Pros: Excellent binlog/relay log delta compensation, minimizing transaction loss during unplanned switchovers.

Cons: Requires passwordless root SSH access between servers, creating significant security vulnerabilities. The project is no longer actively maintained. Additionally, it lacks a web interface and becomes cumbersome when managing multi-tier replication topologies (intermediate masters).

3. Orchestrator

Pros:

  • No SSH access needed—only a dedicated MySQL service account with basic replication privileges.
  • An intuitive Web UI with real-time topology visualization. Replicas can be repointed during maintenance via simple drag-and-drop actions.
  • Holistic failure detection: Orchestrator declares a master dead only when all replicas confirm losing upstream connectivity, avoiding false alarms triggered by transient local network partitions.
  • Native support for MySQL GTID and Pseudo-GTID.

Cons: Orchestrator focuses solely on topology state and node promotion. Routing application traffic to the new master still requires integration with ProxySQL, Consul, or custom VIP takeover scripts.

When Should You Switch to Orchestrator?

If you operate a MySQL cluster with 1 Master and 3 or more Read Replicas processing thousands of queries per second, the limitations of Keepalived become evident quickly. In mission-critical environments where you might otherwise look at deploying MySQL NDB Cluster, Orchestrator delivers the ideal balance: enabling single-click manual topology modifications during routine maintenance while providing robust, automated recovery without split-brain risks.

Step-by-Step Orchestrator Implementation Guide

Step 1: Create the MySQL User for Orchestrator

Create a dedicated monitoring user on all cluster nodes (both Primary and Replicas):

-- Run this on the Master and all Replicas
CREATE USER 'orc_client'@'192.168.1.%' IDENTIFIED BY 'MatKhauSieuKho123!';
GRANT SUPER, PROCESS, REPLICATION SLAVE, RELOAD ON *.* TO 'orc_client'@'192.168.1.%';
GRANT SELECT ON performance_schema.* TO 'orc_client'@'192.168.1.%';
FLUSH PRIVILEGES;

Step 2: Prepare the Backend Database for Orchestrator

Orchestrator requires its own backend store to maintain topology state and metadata. While SQLite is suitable for testing, a dedicated MySQL instance is recommended for production environments:

CREATE DATABASE IF NOT EXISTS orchestrator;
CREATE USER 'orc_server'@'127.0.0.1' IDENTIFIED BY 'BackendPassOrc456!';
GRANT ALL PRIVILEGES ON orchestrator.* TO 'orc_server'@'127.0.0.1';
FLUSH PRIVILEGES;

Step 3: Install and Configure the Orchestrator Service

Download and install the package on your management host (Ubuntu/Debian):

curl -LO https://github.com/openark/orchestrator/releases/download/v3.2.6/orchestrator_3.2.6_amd64.deb
sudo dpkg -i orchestrator_3.2.6_amd64.deb

Configure /etc/orchestrator.conf.json with the following essential parameters:

{
  "Debug": false,
  "EnableSyslog": true,
  "ListenAddress": ":3000",
  "MySQLOrchestratorHost": "127.0.0.1",
  "MySQLOrchestratorPort": 3306,
  "MySQLOrchestratorDatabase": "orchestrator",
  "MySQLOrchestratorUser": "orc_server",
  "MySQLOrchestratorPassword": "BackendPassOrc456!",
  "MySQLTopologyUser": "orc_client",
  "MySQLTopologyPassword": "MatKhauSieuKho123!",
  "DiscoverByShowSlaveHosts": true,
  "InstancePollSeconds": 5,
  "RecoveryPeriodBlockSeconds": 300,
  "Processes": {
    "PreGracefulTakeoverProcesses": [],
    "PostMasterFailoverProcesses": [
      "/usr/local/bin/notify_failover.sh {failureType} {failedHost} {successorHost}"
    ]
  }
}

Start and enable the service:

sudo systemctl enable --now orchestrator
sudo systemctl status orchestrator

Step 4: Discover the Replication Topology

Access the Web UI at http://<ORCHESTRATOR_IP>:3000. To initiate discovery, enter the master node IP under Clusters > Discover or run the CLI command:

orchestrator -c discover -i 192.168.1.10:3306

Orchestrator will query 192.168.1.10, run SHOW REPLICA HOSTS, and automatically construct an interactive tree map of your entire topology. This topology visibility is especially helpful when monitoring cluster health, tracking configurations like delayed replication in MySQL, or managing downstream slaves with specific MySQL replication filters.

Step 5: Enable and Validate Automated Failover

To configure Orchestrator to automatically respond to master outages, enable these three flags in /etc/orchestrator.conf.json:

{
  "ApplyMySQLPromotionAfterMasterFailover": true,
  "FailMasterPromotionIfSQLThreadNotUpToDate": true,
  "AutoMasterRecovery": true
}

Reload the configuration: sudo systemctl restart orchestrator.

Now, simulate an outage by stopping the MySQL service on the Master (192.168.1.10):

# Simulate sudden Master crash
sudo systemctl stop mysql

Observe the recovery sequence in real time:

sudo journalctl -u orchestrator -f

Orchestrator systematically executes the following workflow:

  1. Detects loss of direct connectivity to 192.168.1.10 from the Orchestrator daemon.
  2. Queries all downstream replicas. Observing broken IO threads across all child nodes, Orchestrator confirms a consensus DeadMaster state.
  3. Evaluates Executed_Gtid_Set across all candidate replicas and identifies the most up-to-date node (e.g., 192.168.1.11).
  4. Issues promotion commands on 192.168.1.11: stops replication and enables write operations (SET GLOBAL read_only = 0).
  5. Repoints the remaining replicas to synchronize from the newly promoted master at 192.168.1.11.
  6. Triggers the PostMasterFailoverProcesses hook to update routing rules in ProxySQL or adjust DNS records.

The entire failover workflow completes in roughly 10 to 15 seconds. Write availability is restored seamlessly without requiring manual midnight intervention, while preventing issues like replication data drift between cluster nodes.

Share: