# Rate Limiting at the Gateway — API Gateway and Service Mesh

Source: https://www.geekswithgeeks.com/en/api-gateway-service-mesh/s-rate-limit

> Protect backends and be fair to clients with token-bucket style limits.

## Rate plus burst

Gateways rate-limit per key, user or IP. A **token bucket** allows a sustained rate plus a short **burst**: the bucket holds up to `burst` tokens, refills at `rate` per second, and each request spends one. In nginx, `limit_req_zone ... rate=5r/s` defines the rate and `limit_req ... burst=5 nodelay` the burst; excess requests are rejected with the status set by `limit_req_status` (use **429**). Return `Retry-After` and rate-limit headers so well-behaved clients back off. In a cluster of gateway instances, keep counters in a shared store (for example Redis) or accept that each instance enforces its own share.

## Rate limit config

5 requests per second per client IP, with a burst of 5 extra requests allowed immediately. Excerpt of the gateway config.

```nginx
limit_req_zone $binary_remote_addr zone=perip:1m rate=5r/s;
limit_req_status 429;

location /public/ { limit_req zone=perip burst=5 nodelay; proxy_pass http://orders_stable/; }
```

## 20 rapid requests, run

I ran this against a real nginx 1.27 gateway in Docker, with small Node.js services as upstreams (full setup in the case study). Sending 20 requests at once, the first 6 succeed (the 1 allowed by the rate plus the burst of 5) and the other 14 are rejected with `429`.

```bash
# 20 concurrent calls to /public/x, tallying status codes
```

Output:

```
{"200":6,"429":14}
```

## A token bucket model, run

I ran this plain-Python model. A bucket of capacity 5 refilling at 5 per second allows 5 of 20 instant requests, and again 5 of 20 one second later, once refilled.

```python
import hashlib, bisect, math, random

class Bucket:
    def __init__(self, cap, rate): self.cap, self.rate, self.tokens, self.t = cap, rate, cap, 0.0
    def allow(self, now):
        self.tokens = min(self.cap, self.tokens + (now - self.t) * self.rate); self.t = now
        if self.tokens >= 1: self.tokens -= 1; return True
        return False
bk = Bucket(5, 5.0)
print(sum(bk.allow(0.0) for _ in range(20)), "allowed of 20 at t=0;", sum(bk.allow(1.0) for _ in range(20)), "of 20 one second later")

```

Output:

```
5 allowed of 20 at t=0; 5 of 20 one second later
```

**Quiz:** What does the burst setting control?

- [x] How many requests above the steady rate can be accepted instantly
- [ ] The size of response bodies
- [ ] The TLS version
- [ ] The log format

*Answer:* How many requests above the steady rate can be accepted instantly. Burst absorbs short spikes while the rate limits the long-run average.
