Lesson 10 / 26
Rate Limiting at the Gateway
Protect backends and be fair to clients with token-bucket style limits.
Rate plus burst
Gateways rate-limit per key, user or IP. A token bucket allows a sustained rate plus a short burst: the bucket holds up to burst tokens, refills at rate per second, and each request spends one. In nginx, limit_req_zone ... rate=5r/s defines the rate and limit_req ... burst=5 nodelay the burst; excess requests are rejected with the status set by limit_req_status (use 429). Return Retry-After and rate-limit headers so well-behaved clients back off. In a cluster of gateway instances, keep counters in a shared store (for example Redis) or accept that each instance enforces its own share.
Rate limit config
5 requests per second per client IP, with a burst of 5 extra requests allowed immediately. Excerpt of the gateway config.
limit_req_zone $binary_remote_addr zone=perip:1m rate=5r/s;
limit_req_status 429;
location /public/ { limit_req zone=perip burst=5 nodelay; proxy_pass http://orders_stable/; }20 rapid requests, run
I ran this against a real nginx 1.27 gateway in Docker, with small Node.js services as upstreams (full setup in the case study). Sending 20 requests at once, the first 6 succeed (the 1 allowed by the rate plus the burst of 5) and the other 14 are rejected with 429.
# 20 concurrent calls to /public/x, tallying status codes
Output:
{"200":6,"429":14}A token bucket model, run
I ran this plain-Python model. A bucket of capacity 5 refilling at 5 per second allows 5 of 20 instant requests, and again 5 of 20 one second later, once refilled.
import hashlib, bisect, math, random
class Bucket:
def __init__(self, cap, rate): self.cap, self.rate, self.tokens, self.t = cap, rate, cap, 0.0
def allow(self, now):
self.tokens = min(self.cap, self.tokens + (now - self.t) * self.rate); self.t = now
if self.tokens >= 1: self.tokens -= 1; return True
return False
bk = Bucket(5, 5.0)
print(sum(bk.allow(0.0) for _ in range(20)), "allowed of 20 at t=0;", sum(bk.allow(1.0) for _ in range(20)), "of 20 one second later")
Output:
5 allowed of 20 at t=0; 5 of 20 one second later
Quick check: What does the burst setting control?
- How many requests above the steady rate can be accepted instantly
- The size of response bodies
- The TLS version
- The log format
Answer
How many requests above the steady rate can be accepted instantly — Burst absorbs short spikes while the rate limits the long-run average.