Lesson 17 / 25
Retries and Exponential Backoff
Retry transient errors with growing delays and give up on permanent ones.
Retry the right errors
Network blips, rate limits (HTTP 429) and server errors (5xx) are often transient, so retrying after a wait can succeed. Bad requests (400) or permission errors (403) are permanent: retrying only wastes money. Use exponential backoff (1s, 2s, 4s, 8s…) with a cap, add a little random jitter so many clients do not retry in lockstep, and limit the number of tries.
Delay schedule
This ran as shown: delays double until they reach the cap of 8 seconds. Add random.uniform(0, 0.5) to each in real code.
def delays(base=1.0, cap=8.0, tries=5):
return [min(cap, base * 2 ** i) for i in range(tries)]
print(delays())
Output:
[1.0, 2.0, 4.0, 8.0, 8.0]
Respect Retry-After
When a response includes a Retry-After header, wait at least that long instead of using your own schedule. Many SDKs already retry transient errors for you; check before adding your own layer.
Quick check: Which error is usually worth retrying after a wait?
- 400 Bad Request
- 403 Forbidden
- 429 Too Many Requests
- 404 Not Found for a wrong ID
Answer
429 Too Many Requests — Rate limits are temporary; the others will fail the same way until the request changes.