Lesson 11 / 27
Rate Limiting and Quotas
Protect the service and be fair to clients with limits and clear signals.
Limits are a feature
Without limits, one buggy loop or abusive client can take the service down for everyone. Apply rate limits per API key, user or IP (for example 100 requests per minute) and sometimes quotas (per day or month), often with a token-bucket or sliding-window algorithm. Tell clients where they stand: common headers are RateLimit-Limit, RateLimit-Remaining and RateLimit-Reset (an IETF draft standard, with many services still using X-RateLimit-* names), and when blocked return 429 Too Many Requests with Retry-After (seconds or a date). Well-behaved clients then back off. Stricter limits fit expensive endpoints (search, exports) and login attempts.
Hitting the limit, run
I ran this against a small real HTTP API built with only the Python standard library (full code in the case study). The /limited route allows 3 requests. The headers count down remaining requests (2, 1, 0), then the 4th and 5th calls get 429 with Retry-After: 30. (Output columns: status, remaining, retry-after.)
# uses srv, base and call() from the runnable demo in the case study
for _ in range(5):
s, h, b = call("GET", "/limited"); print(s, h.get("X-RateLimit-Remaining"), h.get("Retry-After"))
srv.shutdown()
Output:
200 2 None 200 1 None 200 0 None 429 0 30 429 0 30
Client-side backoff, run
I ran this plain-Python example. A client should honour Retry-After when present, and otherwise back off exponentially up to a cap: 1, 2, 4, 8, 16, then 30 seconds. With Retry-After: 30 it waits exactly 30.
import hmac, hashlib, json, time, bisect
def next_delay(attempt, retry_after=None, base=1.0, cap=30.0):
return float(retry_after) if retry_after is not None else min(cap, base * 2 ** attempt)
print([next_delay(a) for a in range(6)], next_delay(2, "30"))
Output:
[1.0, 2.0, 4.0, 8.0, 16.0, 30.0] 30.0
Quick check: Which status and header tell a client to wait?
- 301 with Location
- 200 with Content-Length
- 404 with ETag
- 429 with Retry-After
Answer
429 with Retry-After — 429 says "too many requests" and Retry-After says when to try again.