# Capacity Planning With Little's Law — LLM Engineering Foundations

Source: https://www.geekswithgeeks.com/en/llm-engineering/p-capacity

> Estimate concurrency and servers from request rate and latency.

## Requests in flight = rate x time

If you serve your own model, or want to know your provider limits in advance, simple arithmetic goes a long way. **Little's law** says the average number of requests in the system equals the **arrival rate** times the **average time each spends in the system**. LLM calls take seconds, so concurrency grows quickly: 50 requests per second at 4 seconds each means about 200 requests in flight at once. Divide by how many concurrent requests one server (or one rate-limited key) can handle, and **leave headroom** (for example run at 70% of capacity) for bursts and failures. This also shows that **latency and capacity are linked**: halving average latency halves the needed capacity, and a slowdown (longer outputs, a slower model) multiplies it. Peak traffic matters more than the average, so plan from the busiest hour.

## Servers needed from rate and latency, run

I ran this with plain Python 3 (standard library only). At 5 requests/s and 4 s each, 20 requests are in flight and 2 servers of 16 slots suffice at 70% load. At 50 requests/s there are 200 in flight and 18 servers are needed; if latency triples to 12 s, 54 servers are needed; at 200 requests/s and 4 s, 72 servers. The 16-slot servers are an assumed example.

```python
# Little's law: concurrent requests in flight = arrival rate x average time in the system.
def servers_needed(rps, latency_s, per_server_concurrency, headroom=0.7):
    in_flight = rps * latency_s
    return in_flight, -(-in_flight // (per_server_concurrency * headroom))

for rps, lat in ((5, 4.0), (50, 4.0), (50, 12.0), (200, 4.0)):
    inflight, servers = servers_needed(rps, lat, per_server_concurrency=16)
    print(f"{rps:4d} req/s x {lat:4.1f} s -> {inflight:6.0f} requests in flight -> {int(servers)} servers of 16 slots at 70% target load")

```

Output:

```
   5 req/s x  4.0 s ->     20 requests in flight -> 2 servers of 16 slots at 70% target load
  50 req/s x  4.0 s ->    200 requests in flight -> 18 servers of 16 slots at 70% target load
  50 req/s x 12.0 s ->    600 requests in flight -> 54 servers of 16 slots at 70% target load
 200 req/s x  4.0 s ->    800 requests in flight -> 72 servers of 16 slots at 70% target load
```

## Plan from the busiest hour

Peak traffic, not the daily average, decides how much capacity you need.

**Quiz:** If average latency doubles at the same request rate, what happens to requests in flight?

- [ ] They stay the same
- [ ] They halve
- [x] They double
- [ ] They become zero

*Answer:* They double. In-flight requests = rate x latency, so they scale with latency.
