# Retries, Idempotency and Retry Storms — API Gateway and Service Mesh

Source: https://www.geekswithgeeks.com/en/api-gateway-service-mesh/z-retries

> Retry only safe, transient failures, with backoff and budgets, and avoid multiplying load.

## Helpful or harmful

Retrying fixes transient problems (a pod restarting, a dropped connection) but, done carelessly, makes outages worse. Rules: retry only **idempotent** requests (GET, PUT, DELETE, or POST with an idempotency key) and only on **safe failure types** (connection errors, timeouts, `502/503/504`), never on `400/401/404`. Use **exponential backoff with jitter** and a small **retry budget** (for example, retries may add at most 10% extra load). And remember **amplification**: if each of 3 layers retries 2 extra times, one user request can become 27 calls to the deepest service. Retry at **one layer** (often the gateway or the mesh) rather than at every layer.

## Retry on 503 (config)

`resilient` lists a broken upstream (`flaky`, always 503) and a healthy one. `proxy_next_upstream` retries the next server on errors, timeouts and 503. Excerpt of the gateway config.

```nginx
upstream resilient { server flaky:8080; server orders-v1:8080; }

location /resilient/ { proxy_next_upstream error timeout http_503; proxy_pass http://resilient/; }
```

## Every request succeeds despite a broken server, run

I ran this against a real nginx 1.27 gateway in Docker, with small Node.js services as upstreams (full setup in the case study). Four calls all return `200` from `orders-v1`: whenever nginx picked the broken `flaky` server first, it got a 503 and transparently retried on the healthy one. The client never saw the failure.

```bash
# four calls to /resilient/x: status and the service that answered
```

Output:

```
200 orders-v1, 200 orders-v1, 200 orders-v1, 200 orders-v1
```

## Amplification and backoff, run

I ran this plain-Python model. With 2 retries at every layer, 1, 2, 3 and 4 layers of retries produce 3, 9, 27 and 81 calls. The second line is the retry-budget idea: retries capped at 10% extra means about 1.1 times the load. The backoff list shows waits doubling from 0.1 s up to a 5 s cap.

```python
import hashlib, bisect, math, random

def calls(layers, retries): return (1 + retries) ** layers
print([calls(l, 2) for l in (1, 2, 3, 4)])
def with_budget(rate, budget=0.1): return round(1 + budget, 2)
print(with_budget(100))

def backoff(attempt, base=0.1, cap=5.0): return min(cap, base * 2 ** attempt)
print([backoff(a) for a in range(8)])

```

Output:

```
[3, 9, 27, 81]
1.1
[0.1, 0.2, 0.4, 0.8, 1.6, 3.2, 5.0, 5.0]
```

**Quiz:** Why is retrying at every layer dangerous?

- [ ] Retries are not allowed in HTTP
- [x] Retries multiply, turning one request into many and overloading the failing service
- [ ] Retries make responses smaller
- [ ] It only affects logging

*Answer:* Retries multiply, turning one request into many and overloading the failing service. Layered retries amplify load exponentially during an outage.
