# Retries and Exponential Backoff — Agent Loops, Stop Conditions and Token Budgets

Source: https://www.geekswithgeeks.com/en/agent-loops/fail-retries-backoff

> Retry transient errors with growing delays and give up on permanent ones.

## Retry the right errors

Network blips, rate limits (HTTP 429) and server errors (5xx) are often **transient**, so retrying after a wait can succeed. Bad requests (400) or permission errors (403) are **permanent**: retrying only wastes money. Use **exponential backoff** (1s, 2s, 4s, 8s…) with a cap, add a little random jitter so many clients do not retry in lockstep, and limit the number of tries.

## Delay schedule

This ran as shown: delays double until they reach the cap of 8 seconds. Add `random.uniform(0, 0.5)` to each in real code.

```python
def delays(base=1.0, cap=8.0, tries=5):
    return [min(cap, base * 2 ** i) for i in range(tries)]

print(delays())
```

Output:

```
[1.0, 2.0, 4.0, 8.0, 8.0]
```

## Respect Retry-After

When a response includes a `Retry-After` header, wait at least that long instead of using your own schedule. Many SDKs already retry transient errors for you; check before adding your own layer.

**Quiz:** Which error is usually worth retrying after a wait?

- [ ] 400 Bad Request
- [ ] 403 Forbidden
- [x] 429 Too Many Requests
- [ ] 404 Not Found for a wrong ID

*Answer:* 429 Too Many Requests. Rate limits are temporary; the others will fail the same way until the request changes.
