Skip to content

Retry Logic & Errors

LangGraph workflows often rely on external tools (e.g., APIs, databases) that can fail due to network issues, rate limits, or invalid inputs. Robust error handling and retry logic are critical to ensure reliability. This section covers strategies for managing tool call failures, including retries, timeouts, and recovery patterns.


Core Concepts

Retry Strategies

Retrying failed tool calls is a common pattern to handle transient errors (e.g., network glitches). LangGraph integrates with libraries like tenacity to implement retries with exponential backoff and jitter.

from tenacity import retry, stop_after_attempt, wait_exponential_jitter

@retry(stop=stop_after_attempt(3), wait=wait_exponential_jitter(0.1, 0.5))
def call_tool_with_retry():
    result = tool_call()
    if result.status == "failed":
        raise Exception("Tool call failed")
    return result

Key parameters: - stop_after_attempt: Maximum retry attempts. - wait_exponential_jitter: Randomized delays to avoid thundering herd problems.


Timeouts and Circuit Breakers

Timeouts prevent workflows from hanging indefinitely. Circuit breakers stop retrying after a certain number of failures, avoiding infinite loops.

from langgraph import tool_call
import asyncio

async def safe_tool_call():
    try:
        result = await tool_call(timeout=5.0)  # 5-second timeout
        return result
    except asyncio.TimeoutError:
        logger.error("Tool call timed out")
        return {"status": "timeout", "message": "Operation exceeded allowed duration"}

Circuit breaker example (stateful):

failure_count = 0
max_failures = 3
cooldown_period = 60  # seconds

def check_circuit_breaker():
    global failure_count
    if failure_count >= max_failures:
        if time.time() - last_failure_time > cooldown_period:
            failure_count = 0
            return False
        return True
    return False


Failure Recovery Patterns

Fallback Responses

Use fallback logic to provide default responses when tool calls fail:

def handle_tool_failure(error):
    if "rate limit" in error:
        return "Error: API rate limit exceeded. Please try again later."
    elif "network" in error:
        return "Error: Temporary network issue. Retrying..."
    else:
        return "Error: Unknown failure. Please check the tool configuration."

State Persistence

Persist workflow state to retry from a known checkpoint. For example, save intermediate results to a database or file system.

def save_checkpoint(state):
    with open("workflow_checkpoint.json", "w") as f:
        json.dump(state, f)

def load_checkpoint():
    try:
        with open("workflow_checkpoint.json", "r") as f:
            return json.load(f)
    except FileNotFoundError:
        return {}

Monitoring and Logging

Structured Logging

Log detailed error metadata for debugging:

import logging

logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)

def log_tool_error(error, tool_name):
    logger.error(f"[Tool: {tool_name}] Error: {error}")
    logger.debug("Stack trace: %s", error.__traceback__)

Metrics Collection

Track retry counts and failure types using Prometheus or custom metrics:

from prometheus_client import Counter

tool_failure_counter = Counter("tool_failure_total", "Total tool failures by type")

def record_tool_failure(error_type):
    tool_failure_counter.labels(error_type).inc()

Diagram: Retry and Recovery Flow

graph TD
    A[Tool Call] --> B{Success?}
    B -->|Yes| C[Return Result]
    B -->|No| D[Check Retry Policy]
    D -->|Retry| E[Retry Tool Call]
    D -->|No Retry| F[Log Failure]
    F --> G[Trigger Recovery (Fallback/Checkpoint)]

Key takeaways

  • Use exponential backoff with jitter for retries to avoid cascading failures.
  • Combine timeouts with circuit breakers to prevent infinite loops.
  • Implement fallback responses and state persistence for graceful degradation.
  • Log structured error metadata and track metrics for post-mortem analysis.
  • Prioritize idempotency in tool calls to safely retry without side effects.