Skip to content

Resilience Tests

Designing Resilience Tests

Resilience testing is the process of validating a system’s ability to withstand failures, maintain availability, and recover gracefully under stress. By designing targeted chaos experiments, you can test hypotheses about failover mechanisms, redundancy configurations, and system robustness. This section explains how to craft testable hypotheses, select appropriate chaos experiments, and validate resilience outcomes using tools like LitmusChaos and Gremlin.


Defining Testable Hypotheses

A hypothesis in resilience testing is a statement about how a system should behave under specific failure scenarios. It should be specific, measurable, and falsifiable. For example:
- "If the primary database node fails, the system will automatically fail over to a secondary node within 5 seconds."
- "A network partition between two microservices will not cause a cascade failure if circuit breakers are properly configured."

Example: Hypothesis for Database Failover

# Hypothesis: "Database failover will occur within 5 seconds if the primary node is terminated."
# LitmusChaos experiment to test this
litmus create experiment --name db-failover-test \
  --namespace default \
  --chaos-type node-failure \
  --node-name db-primary \
  --timeout 5s \
  --output "failover-verification"

Key Considerations

  • Scope: Focus on critical components (e.g., databases, load balancers, inter-service communication).
  • Metrics: Define success criteria (e.g., latency thresholds, error rates, recovery time).
  • Repeatability: Ensure experiments can be reproduced for consistent results.

Selecting Chaos Experiments

Choose experiments that simulate real-world failure modes relevant to your system. Common chaos types include:
- Node/Service Failure: Simulate hardware or process crashes.
- Network Partitioning: Test isolation between microservices or regions.
- Resource Exhaustion: Stress CPU, memory, or disk I/O.
- Latency Injection: Mimic slow network or backend responses.

Example: Network Partition Experiment

# Gremlin experiment to simulate a network partition between two services
apiVersion: gremlin.gremlin.sh/v1
kind: Experiment
metadata:
  name: network-partition-test
spec:
  duration: 30s
  workload:
    type: network-partition
    parameters:
      target: "service-a"
      peer: "service-b"
      partition: true
  metrics:
    - name: "latency"
      threshold: 500ms
      severity: warning

Diagram: Chaos Experiment Workflow

[ Hypothesis ] --> [ Chaos Experiment ] --> [ Monitoring ] --> [ Result Analysis ] --> [ Validation/Iteration ]

Validating Failover Mechanisms

After running an experiment, validate whether the system’s failover mechanisms work as intended. Use tools like Prometheus and Grafana to monitor:
- Latency spikes during failures.
- Error rates (e.g., HTTP 5xx errors).
- Recovery time (e.g., time to restore normal operations).

Example: Grafana Dashboard for Failover Monitoring

-- Prometheus query to detect service downtime
sum by (job) (count_over_time({job="service-a"}[1m])) 

Postmortem Analysis

  • Success: If the system meets SLOs (e.g., <1% error rate), the hypothesis is validated.
  • Failure: If the system degrades beyond acceptable thresholds, refine redundancy configurations or failover logic.

Analyzing Results and Iterating

Use the error budget (the amount of downtime allowed under SLOs) to evaluate outcomes. For example:
- If a test causes 2% downtime (within the error budget), the system is resilient.
- If it exceeds the budget, prioritize fixes to reduce risk.

Example: Error Budget Calculation

# Calculate error budget usage
error_budget = 0.05  # 5% of SLO
actual_downtime = 0.02  # 2% actual downtime
budget_usage = (actual_downtime / error_budget) * 100
print(f"Error budget usage: {budget_usage:.2f}%")

Key takeaways

  • Hypotheses must be specific and measurable to guide effective testing.
  • Select chaos experiments that mirror real-world failure scenarios (e.g., node failures, network partitions).
  • Validate failover mechanisms using metrics and observability tools like Prometheus and Grafana.
  • Iterate based on results to refine resilience strategies and protect against SLO violations.