Lesson 12 / 26

Timeouts and Latency Budgets

Set explicit timeouts at every hop and budget latency across the chain.

No timeout means waiting forever

A slow dependency is often worse than a dead one, because requests pile up holding connections and threads until the caller also fails: a cascading failure. Set explicit timeouts (connect, read) at the gateway and every client, and make them shorter as you go deeper: if the user-facing budget is 2 seconds, the gateway might allow 1.8 s, a service 1.5 s and its database call 1 s, so inner layers give up before outer ones time out. Include the total in a latency budget: add up the typical time of every hop and see where the time goes.

Fail fast, recover, spread load

Distributed systems fail partly and slowly; proxies add time limits, careful retries, circuit breakers and smart balancing.

Four tools: timeout, retry, breaker, balance.
Figure 4.1 — Timeout, retry, breaker and balance.

A read timeout in the gateway (config)

The slow upstream takes 3 seconds; the gateway gives up after 1 second. Excerpt of the gateway config.

location /slow/ { proxy_read_timeout 1s; proxy_pass http://slow_up/; }

The gateway cuts it off at 1 second, run

I ran this against a real nginx 1.27 gateway in Docker, with small Node.js services as upstreams (full setup in the case study). The request returns 504 Gateway Timeout after about 1 second instead of waiting the full 3 seconds.

GET /slow/x

Output:

504 1s

A latency budget, run

I ran this plain-Python model. Four hops add up to 113 ms, and the slowest (payments, 60 ms) is 53% of the total, so that is where optimisation helps most.

import hashlib, bisect, math, random

hops = {"gateway": 5, "auth": 8, "orders": 40, "payments": 60}
print(sum(hops.values()), "ms total;", round(100 * hops["payments"] / sum(hops.values())), "% in the slowest hop")

Output:

113 ms total; 53 % in the slowest hop

Quick check: Why should inner timeouts be shorter than outer ones?

  • Shorter is always more accurate
  • Inner calls give up and report before the outer caller times out
  • It saves bandwidth only
  • It is required by HTTP
Answer

Inner calls give up and report before the outer caller times out — Otherwise outer layers time out while inner work continues uselessly.