# Timeouts and Latency Budgets — API Gateway and Service Mesh

Source: https://www.geekswithgeeks.com/en/api-gateway-service-mesh/z-timeouts

> Set explicit timeouts at every hop and budget latency across the chain.

## No timeout means waiting forever

A slow dependency is often worse than a dead one, because requests pile up holding connections and threads until the caller also fails: a **cascading failure**. Set **explicit timeouts** (connect, read) at the gateway and every client, and make them **shorter as you go deeper**: if the user-facing budget is 2 seconds, the gateway might allow 1.8 s, a service 1.5 s and its database call 1 s, so inner layers give up before outer ones time out. Include the total in a **latency budget**: add up the typical time of every hop and see where the time goes.

## Fail fast, recover, spread load

Distributed systems fail partly and slowly; proxies add time limits, careful retries, circuit breakers and smart balancing.

![Four tools: timeout, retry, breaker, balance.](assets/figures/api-gateway-service-mesh/section-4-map.svg) — Figure 4.1 — Timeout, retry, breaker and balance.

## A read timeout in the gateway (config)

The `slow` upstream takes 3 seconds; the gateway gives up after 1 second. Excerpt of the gateway config.

```nginx
location /slow/ { proxy_read_timeout 1s; proxy_pass http://slow_up/; }
```

## The gateway cuts it off at 1 second, run

I ran this against a real nginx 1.27 gateway in Docker, with small Node.js services as upstreams (full setup in the case study). The request returns `504 Gateway Timeout` after about 1 second instead of waiting the full 3 seconds.

```bash
GET /slow/x
```

Output:

```
504 1s
```

## A latency budget, run

I ran this plain-Python model. Four hops add up to 113 ms, and the slowest (payments, 60 ms) is 53% of the total, so that is where optimisation helps most.

```python
import hashlib, bisect, math, random

hops = {"gateway": 5, "auth": 8, "orders": 40, "payments": 60}
print(sum(hops.values()), "ms total;", round(100 * hops["payments"] / sum(hops.values())), "% in the slowest hop")

```

Output:

```
113 ms total; 53 % in the slowest hop
```

**Quiz:** Why should inner timeouts be shorter than outer ones?

- [ ] Shorter is always more accurate
- [x] Inner calls give up and report before the outer caller times out
- [ ] It saves bandwidth only
- [ ] It is required by HTTP

*Answer:* Inner calls give up and report before the outer caller times out. Otherwise outer layers time out while inner work continues uselessly.
