Lesson 23 / 29

What Limits Serving Throughput

Understand memory, batching and quantisation at a practical level.

Memory and batching decide how many users fit

When serving your own model, three things dominate. Memory: the weights take roughly parameters times bytes per parameter (a 7-billion-parameter model needs about 14 GB in 16-bit precision), plus a KV cache per active request that grows with context length, so long contexts and many concurrent users compete for GPU memory. Batching: GPUs are efficient when processing many requests together, so serving engines use continuous batching (adding new requests to a running batch) to raise throughput, at some cost in individual latency. Quantisation stores weights in 8 or 4 bits to fit more on a GPU and speed up serving, usually with a small quality loss you should measure on your evaluation set. Also relevant: prompt length (prefill cost grows with input), output length (decoding is sequential), speculative decoding and prefix caching in modern engines. Benchmark with your prompts and concurrency, since vendor numbers use convenient settings.

Memory arithmetic for serving

Rules of thumb for planning; real engines add overhead.

weights memory  ~  parameters x bytes per parameter
   7B model, 16-bit (2 bytes):   ~14 GB      8-bit (1 byte): ~7 GB      4-bit (0.5 byte): ~3.5 GB
   70B model, 16-bit:            ~140 GB     4-bit: ~35 GB

KV cache  ~  2 x layers x kv_heads x head_dim x bytes x tokens x concurrent requests
   -> doubling context length or concurrency roughly doubles KV-cache memory

so:  GPU memory = weights + KV cache (grows with users x context) + activations/overhead
     quantise weights, shorten contexts, or add GPUs when it does not fit

Benchmark with your own prompts

Vendor throughput numbers use convenient prompt lengths and settings.

Quick check: What does quantisation mainly do?

  • Reduces memory (and often cost) by storing weights in fewer bits, with some quality loss to measure
  • Makes the model larger
  • Removes the need for evaluation
  • Encrypts the weights
Answer

Reduces memory (and often cost) by storing weights in fewer bits, with some quality loss to measure — Always re-run your evaluation set on the quantised model.