LitmusChaos Best Practices¶
Chaos engineering with LitmusChaos requires careful planning to ensure experiments are meaningful, repeatable, and aligned with your system's reliability goals. This section outlines best practices for designing experiments with measurable SLI metrics, avoiding false positives, and integrating chaos testing into CI/CD pipelines.
Designing Experiments with Measurable SLI Metrics¶
Align Fault Injection with SLIs/SLOs¶
Every chaos experiment should directly tie to specific SLIs (e.g., latency, error rate, availability) and SLOs. For example, if your SLO guarantees 99.9% availability, design experiments to simulate outages and measure how the system responds.
Example:
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
name: nginx-chaos
spec:
engineName: nginx-chaos
annotation:
chaos.io/namespace: default
workloadSelector:
workloadType: nginx
workloadLabel: app=nginx
chaosServiceAccount: litmus-chaos-sa
experiments:
- name: network-latency
spec:
components:
- component: nginx
metrics:
- name: latency
threshold: "500"
interval: "10s"
severity: "critical"
Use Monitoring Tools for Real-Time Feedback¶
Integrate Prometheus/Grafana to visualize SLI metrics during experiments. For example, a Prometheus query to track HTTP error rates:
Example: Chaos Experiment with SLI Validation¶
Avoiding False Positives¶
Use Canary Releases and Gradual Injection¶
Start with low-intensity faults (e.g., 10% network latency) and gradually increase severity. Deploy chaos experiments to a canary subset of your workload to minimize risk.
Example:
spec:
components:
- component: nginx
parameters:
- name: delay
value: "10s"
# Gradually increase delay in subsequent runs
Validate System Resilience with LitmusChaos¶
Leverage LitmusChaos' built-in validation checks to ensure the system recovers and metrics remain within acceptable bounds.
Example:
Example: Post-Experiment Validation¶
# Check if all validations passed
kubectl get chaosengine nginx-chaos -o jsonpath='{.status.phase}'
# Output: "Completed" if all validations succeeded
Integrating with CI/CD Pipelines¶
Automate Chaos Tests in CI/CD¶
Embed chaos experiments into your pipeline to catch regressions early. For example, run chaos tests in a staging environment before merging to production.
Example GitHub Actions Workflow:
jobs:
chaos-test:
runs-on: ubuntu-latest
steps:
- name: Deploy staging environment
run: kubectl apply -f staging-deployment.yaml
- name: Run chaos experiment
run: kubectl apply -f chaos-experiment.yaml
- name: Validate results
run: |
if [ "$(kubectl get chaosengine nginx-chaos -o jsonpath='{.status.phase}')" == "Completed" ]; then
echo "Chaos test passed"
else
echo "Chaos test failed"
exit 1
fi
Handle Failures Gracefully¶
If an experiment fails, pause the pipeline and investigate. Use postmortems to analyze root causes and improve resilience.
Example:
# Fail the pipeline if chaos test fails
if [ "$CHAOSENGINE_PHASE" != "Completed" ]; then
echo "Chaos test failed: $CHAOSENGINE_PHASE"
exit 1
fi
Key takeaways¶
- Align experiments with SLIs/SLOs to ensure measurable outcomes.
- Use validation checks to avoid false positives and verify system resilience.
- Automate chaos testing in CI/CD pipelines to catch regressions early.
- Monitor metrics in real time using tools like Prometheus and Grafana.
- Prioritize gradual fault injection and canary releases to minimize risk.