Lesson 12 / 31

Retries, Fallbacks and Timeouts

Make chains survive transient failures.

Plan for the failing call

Network calls to model providers fail: rate limits, timeouts, temporary 5xx errors. Every runnable offers .with_retry(...) to retry on errors with a limit on attempts and (optionally) exponential backoff, and .with_fallbacks([...]) to switch to an alternative runnable when it fails (for example a second provider or a cheaper model). Also set timeouts and a maximum number of retries on the model itself, and retry only errors that are safe to repeat; do not blindly retry steps with side effects such as sending an email.

Retry and fallback, run

I ran this offline in a Python virtual environment with langchain-core 1.6.6, langchain-text-splitters 1.1.2 and llama-index-core 0.14.25. No API key or network call is needed because a fake model or a toy embedding stands in for the real one. The flaky function fails twice and succeeds on the third call, and with_retry returns "ok after 3 calls". A runnable that always fails (1/0) falls back to the backup and returns "fallback answer".

from langchain_core.runnables import RunnableLambda

calls = {"n": 0}
def flaky(x):
    calls["n"] += 1
    if calls["n"] < 3:
        raise ValueError("temporary failure")
    return f"ok after {calls['n']} calls"

chain = RunnableLambda(flaky).with_retry(stop_after_attempt=4, wait_exponential_jitter=False)
print(chain.invoke("x"))

backup = RunnableLambda(lambda x: "fallback answer")
broken = RunnableLambda(lambda x: 1 / 0).with_fallbacks([backup])
print(broken.invoke("x"))

Output:

ok after 3 calls
fallback answer

Quick check: When should you NOT blindly retry a step?

  • When it is a pure function
  • When it has side effects such as sending an email
  • When it is a read-only lookup
  • When it is fast
Answer

When it has side effects such as sending an email — Repeating a side-effecting step can duplicate its effect.