Lesson 9 / 25

Rate Limiting and Abuse Control

Protect budgets and availability with per-user limits such as a token bucket.

Limits protect everyone

Every model call costs money and compute. Without limits, one script can run up a large bill or starve other users. Apply per-user and per-IP limits on requests, tokens or spend, cap maximum input and output length, and require sign-in for expensive features. A token bucket allows short bursts but a steady average rate: tokens refill over time, and each request spends one.

A token bucket, run

I ran this with a bucket of 3 tokens refilling 1 per second. Three requests at time 0 pass, the 4th and the one at 0.5 s are refused, then requests at 1.0 s and 2.0 s pass again as the bucket refills.

class Bucket:
    def __init__(self, cap, rate):
        self.cap, self.rate, self.tokens, self.t = cap, rate, cap, 0.0
    def allow(self, now, cost=1):
        self.tokens = min(self.cap, self.tokens + (now - self.t) * self.rate)
        self.t = now
        if self.tokens >= cost:
            self.tokens -= cost
            return True
        return False

b = Bucket(3, 1.0)
print([b.allow(t) for t in (0, 0, 0, 0, 0.5, 1.0, 2.0)])

Output:

[True, True, True, False, False, True, True]

Limit tokens, not just requests

One request with a 100,000-token document costs far more than one with a short question. Meter by tokens (or estimated cost) so limits track real spend.

Quick check: What does a token bucket allow?

  • Only one request ever
  • Unlimited requests
  • Short bursts with a steady average rate
  • No requests at all
Answer

Short bursts with a steady average rate — Capacity permits bursts; the refill rate bounds the long-run average.