Lesson 13 / 26

Retries, Idempotency and Retry Storms

Retry only safe, transient failures, with backoff and budgets, and avoid multiplying load.

Helpful or harmful

Retrying fixes transient problems (a pod restarting, a dropped connection) but, done carelessly, makes outages worse. Rules: retry only idempotent requests (GET, PUT, DELETE, or POST with an idempotency key) and only on safe failure types (connection errors, timeouts, 502/503/504), never on 400/401/404. Use exponential backoff with jitter and a small retry budget (for example, retries may add at most 10% extra load). And remember amplification: if each of 3 layers retries 2 extra times, one user request can become 27 calls to the deepest service. Retry at one layer (often the gateway or the mesh) rather than at every layer.

Retry on 503 (config)

resilient lists a broken upstream (flaky, always 503) and a healthy one. proxy_next_upstream retries the next server on errors, timeouts and 503. Excerpt of the gateway config.

upstream resilient { server flaky:8080; server orders-v1:8080; }

location /resilient/ { proxy_next_upstream error timeout http_503; proxy_pass http://resilient/; }

Every request succeeds despite a broken server, run

I ran this against a real nginx 1.27 gateway in Docker, with small Node.js services as upstreams (full setup in the case study). Four calls all return 200 from orders-v1: whenever nginx picked the broken flaky server first, it got a 503 and transparently retried on the healthy one. The client never saw the failure.

# four calls to /resilient/x: status and the service that answered

Output:

200 orders-v1, 200 orders-v1, 200 orders-v1, 200 orders-v1

Amplification and backoff, run

I ran this plain-Python model. With 2 retries at every layer, 1, 2, 3 and 4 layers of retries produce 3, 9, 27 and 81 calls. The second line is the retry-budget idea: retries capped at 10% extra means about 1.1 times the load. The backoff list shows waits doubling from 0.1 s up to a 5 s cap.

import hashlib, bisect, math, random

def calls(layers, retries): return (1 + retries) ** layers
print([calls(l, 2) for l in (1, 2, 3, 4)])
def with_budget(rate, budget=0.1): return round(1 + budget, 2)
print(with_budget(100))

def backoff(attempt, base=0.1, cap=5.0): return min(cap, base * 2 ** attempt)
print([backoff(a) for a in range(8)])

Output:

[3, 9, 27, 81]
1.1
[0.1, 0.2, 0.4, 0.8, 1.6, 3.2, 5.0, 5.0]

Quick check: Why is retrying at every layer dangerous?

  • Retries are not allowed in HTTP
  • Retries multiply, turning one request into many and overloading the failing service
  • Retries make responses smaller
  • It only affects logging
Answer

Retries multiply, turning one request into many and overloading the failing service — Layered retries amplify load exponentially during an outage.