# What Limits Serving Throughput — LLM Engineering Foundations

Source: https://www.geekswithgeeks.com/en/llm-engineering/s-throughput

> Understand memory, batching and quantisation at a practical level.

## Memory and batching decide how many users fit

When serving your own model, three things dominate. **Memory**: the weights take roughly parameters times bytes per parameter (a 7-billion-parameter model needs about 14 GB in 16-bit precision), plus a **KV cache** per active request that grows with context length, so long contexts and many concurrent users compete for GPU memory. **Batching**: GPUs are efficient when processing many requests together, so serving engines use **continuous batching** (adding new requests to a running batch) to raise throughput, at some cost in individual latency. **Quantisation** stores weights in 8 or 4 bits to fit more on a GPU and speed up serving, usually with a small quality loss you should measure on your evaluation set. Also relevant: **prompt length** (prefill cost grows with input), **output length** (decoding is sequential), speculative decoding and prefix caching in modern engines. Benchmark with **your** prompts and concurrency, since vendor numbers use convenient settings.

## Memory arithmetic for serving

Rules of thumb for planning; real engines add overhead.

```text
weights memory  ~  parameters x bytes per parameter
   7B model, 16-bit (2 bytes):   ~14 GB      8-bit (1 byte): ~7 GB      4-bit (0.5 byte): ~3.5 GB
   70B model, 16-bit:            ~140 GB     4-bit: ~35 GB

KV cache  ~  2 x layers x kv_heads x head_dim x bytes x tokens x concurrent requests
   -> doubling context length or concurrency roughly doubles KV-cache memory

so:  GPU memory = weights + KV cache (grows with users x context) + activations/overhead
     quantise weights, shorten contexts, or add GPUs when it does not fit
```

## Benchmark with your own prompts

Vendor throughput numbers use convenient prompt lengths and settings.

**Quiz:** What does quantisation mainly do?

- [x] Reduces memory (and often cost) by storing weights in fewer bits, with some quality loss to measure
- [ ] Makes the model larger
- [ ] Removes the need for evaluation
- [ ] Encrypts the weights

*Answer:* Reduces memory (and often cost) by storing weights in fewer bits, with some quality loss to measure. Always re-run your evaluation set on the quantised model.
