Lesson 17 / 29

Capacity Planning With Little's Law

Estimate concurrency and servers from request rate and latency.

Requests in flight = rate x time

If you serve your own model, or want to know your provider limits in advance, simple arithmetic goes a long way. Little's law says the average number of requests in the system equals the arrival rate times the average time each spends in the system. LLM calls take seconds, so concurrency grows quickly: 50 requests per second at 4 seconds each means about 200 requests in flight at once. Divide by how many concurrent requests one server (or one rate-limited key) can handle, and leave headroom (for example run at 70% of capacity) for bursts and failures. This also shows that latency and capacity are linked: halving average latency halves the needed capacity, and a slowdown (longer outputs, a slower model) multiplies it. Peak traffic matters more than the average, so plan from the busiest hour.

Servers needed from rate and latency, run

I ran this with plain Python 3 (standard library only). At 5 requests/s and 4 s each, 20 requests are in flight and 2 servers of 16 slots suffice at 70% load. At 50 requests/s there are 200 in flight and 18 servers are needed; if latency triples to 12 s, 54 servers are needed; at 200 requests/s and 4 s, 72 servers. The 16-slot servers are an assumed example.

# Little's law: concurrent requests in flight = arrival rate x average time in the system.
def servers_needed(rps, latency_s, per_server_concurrency, headroom=0.7):
    in_flight = rps * latency_s
    return in_flight, -(-in_flight // (per_server_concurrency * headroom))

for rps, lat in ((5, 4.0), (50, 4.0), (50, 12.0), (200, 4.0)):
    inflight, servers = servers_needed(rps, lat, per_server_concurrency=16)
    print(f"{rps:4d} req/s x {lat:4.1f} s -> {inflight:6.0f} requests in flight -> {int(servers)} servers of 16 slots at 70% target load")

Output:

   5 req/s x  4.0 s ->     20 requests in flight -> 2 servers of 16 slots at 70% target load
  50 req/s x  4.0 s ->    200 requests in flight -> 18 servers of 16 slots at 70% target load
  50 req/s x 12.0 s ->    600 requests in flight -> 54 servers of 16 slots at 70% target load
 200 req/s x  4.0 s ->    800 requests in flight -> 72 servers of 16 slots at 70% target load

Plan from the busiest hour

Peak traffic, not the daily average, decides how much capacity you need.

Quick check: If average latency doubles at the same request rate, what happens to requests in flight?

  • They stay the same
  • They halve
  • They double
  • They become zero
Answer

They double — In-flight requests = rate x latency, so they scale with latency.