Lesson 23 / 28
Unbounded Consumption: Rate Limits and Denial of Wallet
Cap requests, tokens and spend so abuse cannot exhaust service or budget.
Every request costs money
LLM requests are expensive and slow, which makes them attractive targets. An attacker (or a bug, or a runaway agent loop) can send huge prompts, force long outputs, trigger many tool calls, or simply flood the endpoint, causing denial of service or "denial of wallet" (a very large bill). Controls: authenticate users and apply per-user rate limits (requests per minute) and token budgets; cap input size, output tokens and agent steps; set timeouts; apply global and per-feature spend caps with alerts, so an incident ends before the invoice does; use queues and back-pressure; cache repeated results; and require CAPTCHA or stronger identity for anonymous endpoints. Watch for model extraction (many probing queries to copy behaviour) and unusual usage patterns. The example shows a token bucket allowing a burst of 5 then throttling, and a daily cap ending an abusive stream after 123 requests.
A token bucket and a daily spend cap, run
I ran this with plain Python 3 (standard library only). All attacks here are harmless demonstrations on local data, using no real systems. A bucket with a burst of 5 lets the first 5 of 12 quick requests through and refuses the rest until it refills (a request two seconds later is accepted again). With invented prices, a daily cap of $5 ends an abusive stream of large requests after 123 of them, having spent $4.98.
class TokenBucket:
def __init__(self, rate_per_min, burst):
self.rate, self.burst, self.tokens, self.t = rate_per_min / 60, burst, burst, 0.0
def allow(self, now, cost=1):
self.tokens = min(self.burst, self.tokens + (now - self.t) * self.rate); self.t = now
if self.tokens >= cost: self.tokens -= cost; return True
return False
b = TokenBucket(rate_per_min=30, burst=5)
results = [b.allow(t * 0.1) for t in range(12)] # 12 requests in 1.1 seconds
print("allowed:", "".join("Y" if r else "n" for r in results), "->", sum(results), "of 12")
print("two seconds later:", b.allow(3.2))
# Spend cap: "denial of wallet" protection
price_in, price_out = 3.0, 15.0 # example USD per million tokens
daily_cap_usd, spent = 5.0, 0.0
requests = 0
while True:
cost = (6000 * price_in + 1500 * price_out) / 1e6 # a single abusive request: huge prompt + long answer
if spent + cost > daily_cap_usd: break
spent += cost; requests += 1
print(f"abusive requests served before the ${daily_cap_usd:.0f} daily cap: {requests} (spent ${spent:.2f})")
Output:
allowed: YYYYYnnnnnnn -> 5 of 12 two seconds later: True abusive requests served before the $5 daily cap: 123 (spent $4.98)
Cap agent loops
A step limit stops a confused agent from burning budget in a loop.
Quick check: What does a spend cap with alerts protect against?
- A surprise bill from abuse, bugs or runaway loops
- Spelling mistakes
- Slow typing
- Missing semicolons
Answer
A surprise bill from abuse, bugs or runaway loops — Budgets and rate limits bound the damage of abuse.