Advanced Scenarios
Advanced Kubernetes Chaos Scenarios¶
Kubernetes-native applications require rigorous testing of resilience under complex, real-world failure conditions. Advanced chaos scenarios simulate multi-dimensional failures that challenge system robustness, such as pod eviction, storage I/O latency, and network latency injection. These experiments help validate fault tolerance, auto-scaling, and recovery mechanisms in distributed systems.
Pod Eviction Scenario¶
Overview
Pod eviction simulates node resource exhaustion (CPU, memory, or disk) to trigger Kubernetes' eviction mechanisms. This tests how applications handle sudden pod termination and whether recovery processes (e.g., restarts, rescheduling) are effective.
Example Command
gremlin attack run k8s-pod-eviction \
--namespace <namespace> \
--pod-name <pod-name> \
--reason "ResourceExhaustion" \
--duration 30s
ResourceExhaustion event, forcing Kubernetes to terminate the pod.
Diagram Description
A diagram would show:
1. A Kubernetes node running a pod.
2. Gremlin injecting a resource exhaustion event.
3. The kubelet evicting the pod.
4. The kube-scheduler rescheduling the pod to another node.
Key Considerations
- Ensure the pod has a restartPolicy of Always or OnFailure for meaningful results.
- Monitor metrics like kube_pod_eviction_total in Prometheus to track eviction events.
Storage I/O Latency Injection¶
Overview
Storage I/O latency simulates delays in read/write operations to persistent volumes (PVs), testing how applications handle degraded storage performance. This is critical for databases, stateful apps, or workloads relying on fast disk access.
Example Command
gremlin attack run k8s-storage-io-latency \
--namespace <namespace> \
--storage-class <storage-class> \
--latency 500ms \
--duration 60s
Diagram Description
A diagram would show:
1. A pod accessing a PV.
2. Gremlin injecting latency into the storage I/O path.
3. The application experiencing increased query latency or timeouts.
4. The system's response (e.g., retries, fallback mechanisms).
Key Considerations
- Test with storage classes used by critical workloads.
- Use Prometheus metrics like node_disk_read_time_seconds to monitor I/O performance.
Network Latency Injection¶
Overview
Network latency injection simulates slow or unstable network connections between services, testing microservices resilience. This is vital for distributed systems where communication delays can cascade into failures.
Example Command
gremlin attack run k8s-network-latency \
--namespace <namespace> \
--target <service-name> \
--latency 1500ms \
--protocol tcp \
--duration 45s
Diagram Description
A diagram would show:
1. Two pods/services communicating over a network.
2. Gremlin injecting latency into the network path.
3. The receiving service experiencing delayed responses.
4. The system's handling of timeouts or retries.
Key Considerations
- Use tools like tcpdump or Wireshark to validate latency injection.
- Monitor service-level metrics (e.g., http_request_duration_seconds) for performance degradation.
Key takeaways¶
- Pod eviction tests resilience to sudden node failures and rescheduling.
- Storage I/O latency validates stateful workloads under degraded disk performance.
- Network latency ensures microservices handle slow or unstable communication.
- Combine these scenarios with observability tools (Prometheus, Grafana) to analyze system behavior.
- Always validate chaos experiments with rollback strategies and postmortem analysis.