# Retry Policies and Error Handling — LangGraph Agents & Multi-Agent Systems

Source: https://www.geekswithgeeks.com/en/langgraph-agents/h-retry

> Absorb transient failures at the node level.

## Retry the node, not the whole run

External calls (model APIs, databases, web services) fail transiently. You can attach a **`RetryPolicy`** to a node with the maximum number of attempts, initial interval, backoff factor and which exceptions to retry, so a flaky call is repeated without restarting the graph. Retry only errors that are safe to repeat, and be careful with nodes that have side effects (use idempotency keys). For permanent failures, catch the error in the node and write an **error field to state** so a conditional edge can route to a fallback, a human or a clean failure message. Combine with timeouts on every external call.

## A node retried until it works, run

I ran this offline with langgraph 1.2.12 and langchain-core 1.6.6 in a Python virtual environment. No API key or model is needed because plain Python functions stand in for the model, so the output is repeatable. The node raises `ConnectionError` on the first two calls. The `RetryPolicy` repeats it, and the third call succeeds: "ok after 3 calls".

```python
from typing import TypedDict
from langgraph.graph import StateGraph, START, END
from langgraph.types import RetryPolicy

calls = {"n": 0}
class State(TypedDict):
    out: str

def flaky(state):
    calls["n"] += 1
    if calls["n"] < 3:
        raise ConnectionError("temporary")
    return {"out": f"ok after {calls['n']} calls"}

g = StateGraph(State)
g.add_node("flaky", flaky, retry_policy=RetryPolicy(max_attempts=4, initial_interval=0.01, jitter=False))
g.add_edge(START, "flaky"); g.add_edge("flaky", END)
print(g.compile().invoke({"out": ""}))

```

Output:

```
{'out': 'ok after 3 calls'}
```

**Quiz:** For a permanent failure, what is a good pattern?

- [ ] Increase the temperature
- [ ] Retry forever
- [ ] Ignore it silently
- [x] Record an error in state and route to a fallback or human

*Answer:* Record an error in state and route to a fallback or human. Explicit error state makes failures handleable by the graph.
