Skip to content

Chaos Mesh & Litmus

Orchestrating Multi-Step Chaos Experiments with Chaos Mesh and LitmusChaos

Chaos engineering experiments often require complex, multi-step scenarios to simulate real-world failure sequences. Tools like Chaos Mesh and LitmusChaos provide orchestration capabilities to define conditional fault injection, parallel execution, and sequential fault injection across services. This section demonstrates how to structure experiments with these advanced patterns.


🔄 Multi-Step Chaos Orchestration

Multi-step experiments simulate sequential failures, such as first injecting a network delay and then a pod crash if the system fails to recover. Chaos Mesh and LitmusChaos allow you to chain steps using steps in the experiment definition.

Example: Sequential Network Delay → Pod Crash

apiVersion: chaos-mesh.org/v1alpha1
kind: ChaosEngine
metadata:
  name: network-delay-then-pod-crash
spec:
  parallelism: 1
  workloadSelector:
    labelSelector:
      matchLabels:
        app: my-app
  steps:
  - name: network-delay
    spec:
      failureRate: 100
      delay: 30s
      action: network-delay
      namespace: default
  - name: pod-crash
    spec:
      action: pod-crash
      namespace: default
      mode: one
      selector:
        labelSelector:
          matchLabels:
            app: my-app

Execution command:

kubectl apply -f multi-step-experiment.yaml

Post-experiment cleanup:

kubectl delete -f multi-step-experiment.yaml


🧠 Conditional Fault Injection

Conditional execution allows you to apply faults only if certain criteria are met (e.g., a service remains unavailable after a delay). Both tools support this via when clauses or expected outcomes.

Example: Crash Pod Only If Service Is Unavailable

apiVersion: litmuschaos.io/v1alpha1
kind: ChaosExperiment
metadata:
  name: conditional-pod-crash
spec:
  components:
    - name: network-delay
      spec:
        delay: 30s
        namespace: default
        action: network-delay
    - name: pod-crash
      spec:
        action: pod-crash
        namespace: default
        mode: one
        selector:
          labelSelector:
            matchLabels:
              app: my-app
        when:
          - condition: "service-unavailable"
            value: "true"

Key concepts: - Use when to trigger subsequent steps based on observed states (e.g., HTTP 503 responses). - Monitor metrics (e.g., Prometheus) to validate conditions.


🚀 Parallel Fault Execution

Parallel execution simulates simultaneous failures, such as injecting network latency and CPU load across different services. Chaos Mesh and LitmusChaos allow you to define independent steps that run concurrently.

Example: Network Delay + CPU Load in Parallel

apiVersion: chaos-mesh.org/v1alpha1
kind: ChaosEngine
metadata:
  name: parallel-chaos
spec:
  parallelism: 2
  workloadSelector:
    labelSelector:
      matchLabels:
        app: my-app
  steps:
  - name: network-delay
    spec:
      action: network-delay
      namespace: default
      delay: 20s
  - name: cpu-load
    spec:
      action: cpu-load
      namespace: default
      cpu: "80"

Execution command:

kubectl apply -f parallel-experiment.yaml

Post-experiment cleanup:

kubectl delete -f parallel-experiment.yaml


📊 Diagram: Orchestration Workflow

[Start] 
  ↓
[Step 1: Network Delay] → [Check Service Availability]
  ↓
[Conditional: If Unavailable] → [Step 2: Pod Crash]
  ↓
[Parallel Step: CPU Load] 
  ↓
[End]

Key takeaways

  • Multi-step orchestration enables sequential fault injection to simulate cascading failures.
  • Conditional execution allows dynamic fault injection based on system state or metrics.
  • Parallel execution simulates simultaneous failures across services, testing resilience under complex scenarios.
  • Always validate conditions using monitoring tools (e.g., Prometheus + Grafana) and document postmortems for each experiment.