Skip to content

Troubleshooting

High availability clusters require regular maintenance and proactive troubleshooting to ensure reliability. This section covers procedures for managing node downtime, migrating resources, and resolving common cluster failures using Pacemaker and related tools.


Node Maintenance and Downtime

Graceful Node Shutdowns

To perform maintenance on a node, first ensure all critical resources are migrated. Use pcs cluster stop to stop a node gracefully:

pcs cluster stop <node-name>
Verify resource migration with crm_mon -1 or pcs status. After maintenance, restart the node and re-enable cluster services:
pcs cluster start <node-name>
pcs cluster enable <node-name>

Handling Unplanned Downtime

If a node fails unexpectedly, Pacemaker may trigger fencing. Check the cluster status:

pcs status
If fencing is active, disable it temporarily to investigate:
pcs cluster disable <node-name>
After resolving the issue, re-enable the node and rejoin the cluster.


Resource Migration and Failover

Moving Resources Manually

Use pcs resource move to relocate a resource to a specific node:

pcs resource move <resource-id> <target-node>
Verify the move:
pcs status
Ensure dependencies are met by checking resource constraints with pcs constraint show.

Failover Testing

Simulate a node failure to test failover:

pcs cluster stop <node-name> --force
Monitor resource migration with crm_mon -1. After testing, restore the node:
pcs cluster start <node-name>


Common Failure Scenarios and Resolution

Resource Failures

If a resource fails, check its status:

pcs status <resource-id>
Disable the resource temporarily to investigate:
pcs resource disable <resource-id>
After resolving the issue, re-enable it:
pcs resource enable <resource-id>

Communication Issues

Verify network connectivity between nodes:

ping <peer-node-ip>
Check firewall rules to ensure ports (e.g., 3121 for Corosync) are open. Use pcs cluster setup to reconfigure communication if needed.

Fencing Problems

If fencing fails, check the fence agent configuration:

pcs fence show
Test fencing with:
pcs cluster disable <node-name>
pcs cluster enable <node-name>
Review logs for errors:
journalctl -u pacemaker


Key takeaways

  • Plan node maintenance by migrating resources and using pcs cluster stop/start.
  • Manually move resources with pcs resource move and verify with pcs status.
  • Troubleshoot failures by checking resource status, network connectivity, and fencing configurations.
  • Regularly test failover scenarios to validate cluster resilience.
  • Use logs (journalctl, pcs status) to diagnose and resolve persistent issues.