Lesson 9 / 25
Rate Limiting and Abuse Control
Protect budgets and availability with per-user limits such as a token bucket.
Limits protect everyone
Every model call costs money and compute. Without limits, one script can run up a large bill or starve other users. Apply per-user and per-IP limits on requests, tokens or spend, cap maximum input and output length, and require sign-in for expensive features. A token bucket allows short bursts but a steady average rate: tokens refill over time, and each request spends one.
A token bucket, run
I ran this with a bucket of 3 tokens refilling 1 per second. Three requests at time 0 pass, the 4th and the one at 0.5 s are refused, then requests at 1.0 s and 2.0 s pass again as the bucket refills.
class Bucket:
def __init__(self, cap, rate):
self.cap, self.rate, self.tokens, self.t = cap, rate, cap, 0.0
def allow(self, now, cost=1):
self.tokens = min(self.cap, self.tokens + (now - self.t) * self.rate)
self.t = now
if self.tokens >= cost:
self.tokens -= cost
return True
return False
b = Bucket(3, 1.0)
print([b.allow(t) for t in (0, 0, 0, 0, 0.5, 1.0, 2.0)])
Output:
[True, True, True, False, False, True, True]
Limit tokens, not just requests
One request with a 100,000-token document costs far more than one with a short question. Meter by tokens (or estimated cost) so limits track real spend.
Quick check: What does a token bucket allow?
- Only one request ever
- Unlimited requests
- Short bursts with a steady average rate
- No requests at all
Answer
Short bursts with a steady average rate — Capacity permits bursts; the refill rate bounds the long-run average.