Skip to content

Troubleshooting

Kubernetes environments often face challenges that require deep troubleshooting skills. This section explores common issues like node taints, resource contention, and network policies, along with practical strategies to diagnose and resolve them.


Node Taints and Scheduling Conflicts

Node taints prevent pods from scheduling on specific nodes unless they have matching tolerations. Misconfigured taints can lead to pods being evicted or failing to deploy.

Diagnosis:
Use kubectl describe node <node-name> to inspect taints. Look for entries like node-role.kubernetes.io/worker:NoSchedule.

kubectl describe node worker-node-0
Check pod events for taint-related errors:
kubectl describe pod <pod-name> | grep -i taint

Mitigation:
1. Remove or adjust taints:

kubectl taint nodes worker-node-0 node-role.kubernetes.io/worker:NoSchedule-
2. Add tolerations to pods:
spec:
  tolerations:
    - key: "node-role.kubernetes.io/worker"
      operator: "Exists"
      effect: "NoSchedule"

Example: A deployment fails with Taints invalid. Add tolerations to the pod spec or adjust node taints to align with the workload requirements.


Resource Contention and Eviction Pressure

Resource contention occurs when nodes exceed CPU/memory limits, leading to pod evictions. This is often due to misconfigured resource requests/limits or insufficient cluster sizing.

Diagnosis:
Check node resource usage:

kubectl top node
kubectl top pod --all-namespaces
Use kubectl describe pod <pod-name> to see eviction events:
Events: 
  Type     Reason             Age                From     Message
  ----     ------             ----               ----     -------
  Warning  Evicted           2m                kubelet  Running pod was evicted due to memory pressure

Mitigation:
1. Set proper resource limits:

spec:
  containers:
    - name: my-app
      resources:
        requests:
          memory: "256Mi"
          cpu: "500m"
        limits:
          memory: "512Mi"
          cpu: "1000m"
2. Scale horizontally: Use Horizontal Pod Autoscaler (HPA) to adjust replicas based on metrics.
3. Prioritize workloads: Use Kubernetes Quality of Service (QoS) classes to ensure critical pods are not evicted.

Example: A database pod is evicted due to memory pressure. Increase its memory limit and ensure the node has sufficient capacity.


Network Policies and Connectivity Issues

Network policies restrict traffic between pods, and misconfigurations can block essential communication (e.g., between services or external endpoints).

Diagnosis:
List active network policies:

kubectl get networkpolicy -A
Check policy rules:
kubectl describe networkpolicy <policy-name>
Test connectivity using curl or telnet:
curl -v http://<service-name>.<namespace>.svc.cluster.local

Mitigation:
1. Adjust policy rules:

spec:
  ingress:
    - from:
        - podSelector:
            matchLabels:
              app: my-service
      ports:
        - protocol: TCP
          port: 80
2. Use namespace isolation: Apply policies at the namespace level to avoid over-restricting traffic.
3. Verify DNS resolution: Ensure services are reachable via DNS names (e.g., my-service.namespace.svc.cluster.local).

Example: A service fails to connect to a backend. Check the network policy to ensure ingress rules allow traffic from the service's namespace.


Key takeaways

  • Node taints require balancing tolerations and taint configurations to ensure proper scheduling.
  • Resource contention demands careful monitoring and proper resource allocation to prevent evictions.
  • Network policies must be tested rigorously to avoid blocking critical communication paths.
  • Always combine diagnostic commands (kubectl describe, top, curl) with policy reviews to isolate root causes.