Chaos Mesh & Litmus
Orchestrating Multi-Step Chaos Experiments with Chaos Mesh and LitmusChaos¶
Chaos engineering experiments often require complex, multi-step scenarios to simulate real-world failure sequences. Tools like Chaos Mesh and LitmusChaos provide orchestration capabilities to define conditional fault injection, parallel execution, and sequential fault injection across services. This section demonstrates how to structure experiments with these advanced patterns.
🔄 Multi-Step Chaos Orchestration¶
Multi-step experiments simulate sequential failures, such as first injecting a network delay and then a pod crash if the system fails to recover. Chaos Mesh and LitmusChaos allow you to chain steps using steps in the experiment definition.
Example: Sequential Network Delay → Pod Crash¶
apiVersion: chaos-mesh.org/v1alpha1
kind: ChaosEngine
metadata:
name: network-delay-then-pod-crash
spec:
parallelism: 1
workloadSelector:
labelSelector:
matchLabels:
app: my-app
steps:
- name: network-delay
spec:
failureRate: 100
delay: 30s
action: network-delay
namespace: default
- name: pod-crash
spec:
action: pod-crash
namespace: default
mode: one
selector:
labelSelector:
matchLabels:
app: my-app
Execution command:
Post-experiment cleanup:
🧠Conditional Fault Injection¶
Conditional execution allows you to apply faults only if certain criteria are met (e.g., a service remains unavailable after a delay). Both tools support this via when clauses or expected outcomes.
Example: Crash Pod Only If Service Is Unavailable¶
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosExperiment
metadata:
name: conditional-pod-crash
spec:
components:
- name: network-delay
spec:
delay: 30s
namespace: default
action: network-delay
- name: pod-crash
spec:
action: pod-crash
namespace: default
mode: one
selector:
labelSelector:
matchLabels:
app: my-app
when:
- condition: "service-unavailable"
value: "true"
Key concepts:
- Use when to trigger subsequent steps based on observed states (e.g., HTTP 503 responses).
- Monitor metrics (e.g., Prometheus) to validate conditions.
🚀 Parallel Fault Execution¶
Parallel execution simulates simultaneous failures, such as injecting network latency and CPU load across different services. Chaos Mesh and LitmusChaos allow you to define independent steps that run concurrently.
Example: Network Delay + CPU Load in Parallel¶
apiVersion: chaos-mesh.org/v1alpha1
kind: ChaosEngine
metadata:
name: parallel-chaos
spec:
parallelism: 2
workloadSelector:
labelSelector:
matchLabels:
app: my-app
steps:
- name: network-delay
spec:
action: network-delay
namespace: default
delay: 20s
- name: cpu-load
spec:
action: cpu-load
namespace: default
cpu: "80"
Execution command:
Post-experiment cleanup:
📊 Diagram: Orchestration Workflow¶
[Start]
↓
[Step 1: Network Delay] → [Check Service Availability]
↓
[Conditional: If Unavailable] → [Step 2: Pod Crash]
↓
[Parallel Step: CPU Load]
↓
[End]
Key takeaways¶
- Multi-step orchestration enables sequential fault injection to simulate cascading failures.
- Conditional execution allows dynamic fault injection based on system state or metrics.
- Parallel execution simulates simultaneous failures across services, testing resilience under complex scenarios.
- Always validate conditions using monitoring tools (e.g., Prometheus + Grafana) and document postmortems for each experiment.