# Exponential Backoff With Jitter — Claude API / OpenAI API Basics

Source: https://www.geekswithgeeks.com/en/llm-apis/r-backoff

> Retry politely so you do not make an overload worse.

## Wait longer each time, and randomise

If many clients retry immediately after a failure, they hit the service together and make the problem worse (a **retry storm**). **Exponential backoff** doubles the wait after each failed attempt (0.5 s, 1 s, 2 s, 4 s...) up to a cap, and **jitter** randomises each wait so clients spread out. Also: set a **maximum number of attempts**, respect `retry-after` when the server sends it, retry only **idempotent** or safe operations (a plain model call is safe to repeat, but a tool that sends an email is not), and give up with a clear error or a fallback. Retries cost money and time, so cap total elapsed time too.

## Backoff delays, run

I ran this plain-Python (standard library only) example. The ceilings double from 0.5 s up to the 8 s cap. With full jitter each delay is a random number between 0 and its ceiling (a fixed seed makes the run repeatable), so clients do not retry in lockstep.

```python
import random

def backoff_delays(attempts, base=0.5, cap=8.0, seed=1):
    rnd = random.Random(seed)
    out = []
    for n in range(attempts):
        ceiling = min(cap, base * (2 ** n))          # exponential growth, capped
        out.append(round(rnd.uniform(0, ceiling), 2))   # "full jitter": random up to the ceiling
    return out

print("ceilings:", [min(8.0, 0.5 * 2 ** n) for n in range(6)])
print("delays  :", backoff_delays(6))

```

Output:

```
ceilings: [0.5, 1.0, 2.0, 4.0, 8.0, 8.0]
delays  : [0.07, 0.85, 1.53, 1.02, 3.96, 3.6]
```

## Use the SDK retries first

The official SDKs already retry transient errors a few times. Add your own layer only for needs they do not cover.

**Quiz:** What is the purpose of jitter in backoff?

- [ ] Encrypting requests
- [ ] Making retries faster
- [x] Spreading retries out so clients do not all return at once
- [ ] Choosing the model

*Answer:* Spreading retries out so clients do not all return at once. Randomised waits prevent synchronised retry storms.
