Troubleshooting
High availability clusters require regular maintenance and proactive troubleshooting to ensure reliability. This section covers procedures for managing node downtime, migrating resources, and resolving common cluster failures using Pacemaker and related tools.
Node Maintenance and Downtime¶
Graceful Node Shutdowns¶
To perform maintenance on a node, first ensure all critical resources are migrated. Use pcs cluster stop to stop a node gracefully:
crm_mon -1 or pcs status. After maintenance, restart the node and re-enable cluster services:
Handling Unplanned Downtime¶
If a node fails unexpectedly, Pacemaker may trigger fencing. Check the cluster status:
If fencing is active, disable it temporarily to investigate: After resolving the issue, re-enable the node and rejoin the cluster.Resource Migration and Failover¶
Moving Resources Manually¶
Use pcs resource move to relocate a resource to a specific node:
pcs constraint show.
Failover Testing¶
Simulate a node failure to test failover:
Monitor resource migration withcrm_mon -1. After testing, restore the node:
Common Failure Scenarios and Resolution¶
Resource Failures¶
If a resource fails, check its status:
Disable the resource temporarily to investigate: After resolving the issue, re-enable it:Communication Issues¶
Verify network connectivity between nodes:
Check firewall rules to ensure ports (e.g., 3121 for Corosync) are open. Usepcs cluster setup to reconfigure communication if needed.
Fencing Problems¶
If fencing fails, check the fence agent configuration:
Test fencing with: Review logs for errors:Key takeaways¶
- Plan node maintenance by migrating resources and using
pcs cluster stop/start. - Manually move resources with
pcs resource moveand verify withpcs status. - Troubleshoot failures by checking resource status, network connectivity, and fencing configurations.
- Regularly test failover scenarios to validate cluster resilience.
- Use logs (
journalctl,pcs status) to diagnose and resolve persistent issues.