Skip to content

Cluster Alerts

Pacemaker clusters rely on timely alerts to notify administrators of critical events such as node failures, resource downtime, or configuration issues. This section covers methods to configure alerts using systemd, email, and external monitoring tools, ensuring proactive maintenance and rapid response.


Using systemd for Cluster Alerts

Systemd can be leveraged to run custom scripts that monitor cluster health and trigger alerts. Create a systemd service unit to periodically check the cluster status and execute alert actions.

Example: Create a systemd service for cluster health checks

sudo nano /etc/systemd/system/cluster-alert.service
Add the following content:
[Unit]
Description=Cluster Health Alert Service
Timer=cluster-alert.timer

[Service]
Type=simple
ExecStart=/usr/bin/cluster-health-check.sh

Example: Define a timer to run checks every 5 minutes

sudo nano /etc/systemd/system/cluster-alert.timer
Add:
[Unit]
Description=Run cluster health checks every 5 minutes

[Timer]
OnCalendar=*:0/5
Unit=cluster-alert.service

Example: Script to trigger alerts

#!/bin/bash
CLUSTER_STATUS=$(crm_mon -1 | grep "stack is online")
if [[ ! $CLUSTER_STATUS ]]; then
  echo "Cluster is offline!" | mail -s "Pacemaker Alert" [email protected]
fi
Enable and start the service:
sudo systemctl enable --now cluster-alert.timer


Email Alerts with mail or sendmail

Configure the cluster to send email notifications for critical events. Use tools like mail or sendmail to relay alerts.

Example: Script to send email alerts

#!/bin/bash
if ! crm_mon -1 | grep -q "online"; then
  echo "CRITICAL: Cluster is offline!" | mail -s "Pacemaker Alert" [email protected]
fi
Ensure the mail transfer agent (MTA) is installed and configured (e.g., postfix or sendmail). Note that the mail command requires an MTA to be installed and properly configured before use.


External Monitoring Tools

Integrate Pacemaker with external tools like Nagios, Zabbix, or Prometheus for centralized monitoring. Use the Pacemaker REST API or SSH to query cluster status.

Example: Check cluster status via SSH

ssh cluster-node1 "crm_mon -1 | grep 'online'"
if [[ $? -ne 0 ]]; then
  echo "Cluster offline on cluster-node1" | tee /path/to/alert.log
fi

Example: Nagios plugin for cluster health

#!/bin/bash
if crm_mon -1 | grep -q "online"; then
  echo "OK"
  exit 0
else
  echo "CRITICAL: Cluster offline"
  exit 2
fi
Place this script in Nagios' plugin directory and configure checks in nagios.cfg.


Testing and Validation

  1. Simulate node failure:
    pcs cluster stop node1
    
    Note: This command requires a multi-node cluster. For single-node setups, use alternative methods like stopping the cluster service or simulating failures through other means.
  2. Verify alerts are triggered via email, logs, or monitoring tools.
  3. Restore the node and ensure alerts resolve.

Key takeaways

  • Use systemd to run periodic health checks and trigger alerts via scripts.
  • Configure email alerts using mail or sendmail for quick notifications.
  • Integrate with external tools like Nagios or Zabbix for centralized monitoring.
  • Test alerts by simulating failures to ensure reliability.
  • Always validate alert configurations against your environment's specific tools and policies.