1. Real-World Problem: 2 AM Server Crashes and the Frustration of Missing Metrics
At 2 AM on a Saturday, the API system started throwing widespread 504 Gateway Timeout errors. With clients complaining and my manager calling, I rushed to open my laptop, SSH into the server, and ran top, free -m, and df -h. Ironically, CPU usage sat at just 3%, available RAM exceeded 10GB, and disk capacity was well over 40% free. The outage lasted about 5 minutes and then vanished on its own, leaving behind a clean crime scene.
Without historical metrics, it was pure guesswork. Was it a CPU spike from a background cron job? A memory leak triggering the OOM killer? Or disk I/O bottlenecked by database flushing? The honest answer: nobody knows.
2. Root Cause Analysis: Why Traditional CLI Tools Fall Short
Classic CLI utilities like top, vmstat, or iostat only show the server status at the exact second you hit Enter. They are merely instant snapshots. Once the incident passes, all traces disappear.
Moreover, application logs only capture application-level errors (such as DB connection timeouts)—they will not tell you that disk write latency spiked to 500ms at that exact moment. As your infrastructure scales from 1 to 10 or 50 servers, SSHing into individual machines for manual checks becomes impossible. You need an automated system to periodically collect metrics and persist them in a dedicated Time-Series Database (TSDB).
3. Comparing Infrastructure Monitoring Solutions
- SaaS (Datadog, New Relic): Single-command installation with polished dashboards. However, costs run around $15–$23/host/month. As your cluster grows with millions of custom metrics, the monthly bill quickly becomes prohibitive.
- Zabbix / Nagios: Extremely robust for basic up/down availability checks. On the downside, agent configuration is cumbersome, resource-heavy, and time-series data handling is far less flexible than Prometheus.
- Custom Bash Scripts pushing to MySQL/InfluxDB: Quick to write initially, but difficult to scale, prone to formatting errors, and high-maintenance for the underlying storage.
4. The Solution: The Prometheus, Node Exporter, and Grafana Stack
This is the de facto standard in the DevOps world thanks to being lightweight, open-source, and highly scalable:
- Node Exporter: An agent that collects hardware metrics (CPU, RAM, Disk, Network) on each server and exposes an HTTP endpoint at
:9100/metrics. - Prometheus: The central server that periodically pulls (scrapes) metrics every 15 seconds to store them as time-series data.
- Grafana: The visualization frontend that queries Prometheus data to render dashboards and configure alerts.
Step 1: Create System Users and Configure Directory Permissions
Running monitoring processes as the root user poses unnecessary security risks. Create dedicated service users instead:
sudo useradd --no-create-home --shell /bin/false prometheus
sudo useradd --no-create-home --shell /bin/false node_exporter
sudo mkdir -p /etc/prometheus /var/lib/prometheus
sudo chown -R prometheus:prometheus /etc/prometheus /var/lib/prometheus
Step 2: Install and Configure Node Exporter
Download the stable release (e.g., v1.8.2) from GitHub:
cd /tmp
wget https://github.com/prometheus/node_exporter/releases/download/v1.8.2/node_exporter-1.8.2.linux-amd64.tar.gz
tar -xvf node_exporter-1.8.2.linux-amd64.tar.gz
sudo mv node_exporter-1.8.2.linux-amd64/node_exporter /usr/local/bin/
sudo chown node_exporter:node_exporter /usr/local/bin/node_exporter
rm -rf node_exporter-1.8.2.*
Create a systemd service file to manage the process:
sudo nano /etc/systemd/system/node_exporter.service
File contents:
[Unit]
Description=Node Exporter
Wants=network-online.target
After=network-online.target
[Service]
User=node_exporter
Group=node_exporter
Type=simple
ExecStart=/usr/local/bin/node_exporter
Restart=always
RestartSec=5
[Install]
WantedBy=multi-user.target
Start the service and enable it to launch on boot:
sudo systemctl daemon-reload
sudo systemctl enable --now node_exporter
Verify by inspecting sample metrics output:
curl -s http://localhost:9100/metrics | head -n 10
Step 3: Install and Configure Prometheus Server
Download the Prometheus binary package (version 2.54.1):
cd /tmp
wget https://github.com/prometheus/prometheus/releases/download/v2.54.1/prometheus-2.54.1.linux-amd64.tar.gz
tar -xvf prometheus-2.54.1.linux-amd64.tar.gz
cd prometheus-2.54.1.linux-amd64
sudo mv prometheus promtool /usr/local/bin/
sudo chown prometheus:prometheus /usr/local/bin/prometheus /usr/local/bin/promtool
sudo mv consoles console_libraries /etc/prometheus/
sudo chown -R prometheus:prometheus /etc/prometheus/consoles /etc/prometheus/console_libraries
Create the scrape configuration file:
sudo nano /etc/prometheus/prometheus.yml
Configure Prometheus to periodically scrape metrics from itself and Node Exporter:
global:
scrape_interval: 15s
evaluation_interval: 15s
scrape_configs:
- job_name: 'prometheus'
static_configs:
- targets: ['localhost:9090']
- job_name: 'node_exporter'
static_configs:
- targets: ['localhost:9100']
Update file ownership for the configuration file:
sudo chown prometheus:prometheus /etc/prometheus/prometheus.yml
Create a systemd service for Prometheus:
sudo nano /etc/systemd/system/prometheus.service
[Unit]
Description=Prometheus Server
Wants=network-online.target
After=network-online.target
[Service]
User=prometheus
Group=prometheus
Type=simple
ExecStart=/usr/local/bin/prometheus \
--config.file=/etc/prometheus/prometheus.yml \
--storage.tsdb.path=/var/lib/prometheus/ \
--storage.tsdb.retention.time=30d \
--web.console.templates=/etc/prometheus/consoles \
--web.console.libraries=/etc/prometheus/console_libraries
Restart=always
RestartSec=5
[Install]
WantedBy=multi-user.target
Start the service:
sudo systemctl daemon-reload
sudo systemctl enable --now prometheus
Verify in your browser via port 9090: http://IP_SERVER:9090. Navigate to Status > Targets; if both endpoints display a green UP status, the setup succeeded.
Step 4: Install Grafana
Install Grafana from the official Grafana Labs repository to ensure you receive ongoing security patches:
sudo apt update
sudo apt install -y apt-transport-https software-properties-common wget
sudo mkdir -p /etc/apt/keyrings/
wget -q -O - https://apt.grafana.com/gpg.key | gpg --dearmor | sudo tee /etc/apt/keyrings/grafana.gpg > /dev/null
echo "deb [signed-by=/etc/apt/keyrings/grafana.gpg] https://apt.grafana.com stable main" | sudo tee /etc/apt/sources.list.d/grafana.list
sudo apt update
sudo apt install -y grafana
sudo systemctl enable --now grafana-server
Open your browser and navigate to: http://IP_SERVER:3000. The default credentials are admin / admin (Grafana will prompt you to set a new password on first login).
Step 5: Connect Prometheus and Import a Standard Dashboard
- In the Grafana interface, navigate to Connections > Data sources > click Add data source.
- Select Prometheus.
- In the Prometheus server URL field, enter
http://localhost:9090. - Scroll to the bottom and click Save & test to verify the connection.
- To set up a complete monitoring interface without building every chart manually, go to Dashboards > New > Import.
- Enter the most popular community Dashboard ID for Node Exporter:
1860(Node Exporter Full), then click Load. - Select the Prometheus data source configured in the previous step and click Import.
All server resources will now be visualized in real time: per-core CPU utilization, actual cache vs. available RAM, disk I/O read/write latency, and network throughput.

