# Unbounded Consumption: Rate Limits and Denial of Wallet — LLM Application Security

Source: https://www.geekswithgeeks.com/en/llm-security/p-consume

> Cap requests, tokens and spend so abuse cannot exhaust service or budget.

## Every request costs money

LLM requests are expensive and slow, which makes them attractive targets. An attacker (or a bug, or a runaway agent loop) can send huge prompts, force long outputs, trigger many tool calls, or simply flood the endpoint, causing **denial of service** or **"denial of wallet"** (a very large bill). Controls: **authenticate users** and apply **per-user rate limits** (requests per minute) and **token budgets**; cap **input size**, **output tokens** and **agent steps**; set **timeouts**; apply **global and per-feature spend caps** with alerts, so an incident ends before the invoice does; use **queues and back-pressure**; cache repeated results; and require **CAPTCHA or stronger identity** for anonymous endpoints. Watch for **model extraction** (many probing queries to copy behaviour) and unusual usage patterns. The example shows a token bucket allowing a burst of 5 then throttling, and a daily cap ending an abusive stream after 123 requests.

## A token bucket and a daily spend cap, run

I ran this with plain Python 3 (standard library only). All attacks here are harmless demonstrations on local data, using no real systems. A bucket with a burst of 5 lets the first 5 of 12 quick requests through and refuses the rest until it refills (a request two seconds later is accepted again). With invented prices, a daily cap of $5 ends an abusive stream of large requests after 123 of them, having spent $4.98.

```python
class TokenBucket:
    def __init__(self, rate_per_min, burst):
        self.rate, self.burst, self.tokens, self.t = rate_per_min / 60, burst, burst, 0.0
    def allow(self, now, cost=1):
        self.tokens = min(self.burst, self.tokens + (now - self.t) * self.rate); self.t = now
        if self.tokens >= cost: self.tokens -= cost; return True
        return False

b = TokenBucket(rate_per_min=30, burst=5)
results = [b.allow(t * 0.1) for t in range(12)]            # 12 requests in 1.1 seconds
print("allowed:", "".join("Y" if r else "n" for r in results), "->", sum(results), "of 12")
print("two seconds later:", b.allow(3.2))

# Spend cap: "denial of wallet" protection
price_in, price_out = 3.0, 15.0      # example USD per million tokens
daily_cap_usd, spent = 5.0, 0.0
requests = 0
while True:
    cost = (6000 * price_in + 1500 * price_out) / 1e6          # a single abusive request: huge prompt + long answer
    if spent + cost > daily_cap_usd: break
    spent += cost; requests += 1
print(f"abusive requests served before the ${daily_cap_usd:.0f} daily cap: {requests} (spent ${spent:.2f})")

```

Output:

```
allowed: YYYYYnnnnnnn -> 5 of 12
two seconds later: True
abusive requests served before the $5 daily cap: 123 (spent $4.98)
```

## Cap agent loops

A step limit stops a confused agent from burning budget in a loop.

**Quiz:** What does a spend cap with alerts protect against?

- [x] A surprise bill from abuse, bugs or runaway loops
- [ ] Spelling mistakes
- [ ] Slow typing
- [ ] Missing semicolons

*Answer:* A surprise bill from abuse, bugs or runaway loops. Budgets and rate limits bound the damage of abuse.
