Lesson 16 / 27
Exponential Backoff With Jitter
Retry politely so you do not make an overload worse.
Wait longer each time, and randomise
If many clients retry immediately after a failure, they hit the service together and make the problem worse (a retry storm). Exponential backoff doubles the wait after each failed attempt (0.5 s, 1 s, 2 s, 4 s...) up to a cap, and jitter randomises each wait so clients spread out. Also: set a maximum number of attempts, respect retry-after when the server sends it, retry only idempotent or safe operations (a plain model call is safe to repeat, but a tool that sends an email is not), and give up with a clear error or a fallback. Retries cost money and time, so cap total elapsed time too.
Backoff delays, run
I ran this plain-Python (standard library only) example. The ceilings double from 0.5 s up to the 8 s cap. With full jitter each delay is a random number between 0 and its ceiling (a fixed seed makes the run repeatable), so clients do not retry in lockstep.
import random
def backoff_delays(attempts, base=0.5, cap=8.0, seed=1):
rnd = random.Random(seed)
out = []
for n in range(attempts):
ceiling = min(cap, base * (2 ** n)) # exponential growth, capped
out.append(round(rnd.uniform(0, ceiling), 2)) # "full jitter": random up to the ceiling
return out
print("ceilings:", [min(8.0, 0.5 * 2 ** n) for n in range(6)])
print("delays :", backoff_delays(6))
Output:
ceilings: [0.5, 1.0, 2.0, 4.0, 8.0, 8.0] delays : [0.07, 0.85, 1.53, 1.02, 3.96, 3.6]
Use the SDK retries first
The official SDKs already retry transient errors a few times. Add your own layer only for needs they do not cover.
Quick check: What is the purpose of jitter in backoff?
- Encrypting requests
- Making retries faster
- Spreading retries out so clients do not all return at once
- Choosing the model
Answer
Spreading retries out so clients do not all return at once — Randomised waits prevent synchronised retry storms.