Pacemaker Resources
Pacemaker Resource Management is central to ensuring high availability and reliability in clustered environments. It involves defining, organizing, and controlling resources (services, storage, networks) to maintain their availability across cluster nodes. This section explains how Pacemaker manages resources through resource agents, constraints, and failover policies.
Resource Agents¶
Resource agents are the core components that encapsulate the logic for managing specific services or resources. They abstract the details of starting, stopping, and monitoring resources, allowing Pacemaker to interact with them uniformly.
Key Concepts:
- Standard Agents: Provided by clusters like Corosync or OpenAIS (e.g., apache, mysql, ocfs2).
- External Agents: Third-party tools like heartbeat or drbd can also be integrated.
- Agent Metadata: Each agent has metadata (e.g., CIB schema) describing its capabilities (e.g., start, stop, monitor).
Example: Creating a Resource
apache using the systemd agent for the apache2 service.
Managing Resources:
Use pcs or crm_resource to manipulate resources:
pcs resource disable apache # Temporarily stop the resource
pcs resource enable apache # Restart the resource
crm_resource --resource apache --action status # Check resource status
Constraints¶
Constraints define relationships between resources and nodes, ensuring predictable behavior during failures.
Types of Constraints:
1. Location Constraints: Specify where a resource should run.
apache resource prefers to run on node1.
-
Order Constraints: Define dependencies between resources.
The
mysqlresource must start beforeapache. -
Colocation Constraints: Enforce co-location of resources.
The
apacheresource must run on the same node asmysql.
Note: Constraints are critical for avoiding split-brain scenarios and ensuring service dependencies are met.
Failover Policies¶
Failover policies determine how Pacemaker handles resource failures and node outages.
Stonith Policies:
- stonith-symmetric: Nodes can take over resources after a STONITH (shoot node) action.
- stonith-asymmetric: Only the surviving node can take over resources.
Failure Policy:
The failure-policy attribute defines how Pacemaker responds to node failures:
- reconnect: Attempt to reconnect to the failed node.
- stop: Stop resources if the node fails.
Example: Configuring Failover Policies
Key takeaways¶
- Resource agents abstract service management, enabling uniform control across diverse services.
- Constraints (location, order, colocation) ensure predictable resource placement and dependencies.
- Failover policies (STONITH modes, failure handling) dictate how resources are managed during node failures.
- Tools like
pcsandcrm_resourceare essential for configuring and monitoring resources. - Proper resource management minimizes downtime and ensures cluster resilience.