Lesson 19 / 29

Retry Policies and Error Handling

Absorb transient failures at the node level.

Retry the node, not the whole run

External calls (model APIs, databases, web services) fail transiently. You can attach a RetryPolicy to a node with the maximum number of attempts, initial interval, backoff factor and which exceptions to retry, so a flaky call is repeated without restarting the graph. Retry only errors that are safe to repeat, and be careful with nodes that have side effects (use idempotency keys). For permanent failures, catch the error in the node and write an error field to state so a conditional edge can route to a fallback, a human or a clean failure message. Combine with timeouts on every external call.

A node retried until it works, run

I ran this offline with langgraph 1.2.12 and langchain-core 1.6.6 in a Python virtual environment. No API key or model is needed because plain Python functions stand in for the model, so the output is repeatable. The node raises ConnectionError on the first two calls. The RetryPolicy repeats it, and the third call succeeds: "ok after 3 calls".

from typing import TypedDict
from langgraph.graph import StateGraph, START, END
from langgraph.types import RetryPolicy

calls = {"n": 0}
class State(TypedDict):
    out: str

def flaky(state):
    calls["n"] += 1
    if calls["n"] < 3:
        raise ConnectionError("temporary")
    return {"out": f"ok after {calls['n']} calls"}

g = StateGraph(State)
g.add_node("flaky", flaky, retry_policy=RetryPolicy(max_attempts=4, initial_interval=0.01, jitter=False))
g.add_edge(START, "flaky"); g.add_edge("flaky", END)
print(g.compile().invoke({"out": ""}))

Output:

{'out': 'ok after 3 calls'}

Quick check: For a permanent failure, what is a good pattern?

  • Increase the temperature
  • Retry forever
  • Ignore it silently
  • Record an error in state and route to a fallback or human
Answer

Record an error in state and route to a fallback or human — Explicit error state makes failures handleable by the graph.