# Model Size and Quantization — Large Language Models

Source: https://www.geekswithgeeks.com/en/llms/d-quant

> Estimate hardware needs and shrink models.

## Fewer bits per weight

Model memory is roughly **parameters × bytes per parameter**. **Quantization** stores weights with fewer bits (for example 8-bit or 4-bit instead of 16-bit), cutting memory and often speeding inference, with a usually small and task-dependent quality loss. This is what lets a 7B-parameter model run on a laptop or a modest GPU. Other size levers: **distillation** (train a small model to imitate a big one), **mixture-of-experts** (only part of the network is active per token), and choosing a smaller model for the task. Remember to leave memory for the KV cache and activations.

## Size, speed, choice

Quantization, batching and the open-versus-hosted choice decide cost and control.

![Three questions: fit, serve, choose.](assets/figures/llms/section-7-map.svg) — Figure 7.1 — Fit, serve and choose.

## Weights memory by precision

I ran this plain-Python (standard library only) example. The weights of a 70B model need 140 GB in fp16 but 35 GB at 4-bit; a 7B model drops from 14 GB to 3.5 GB. These are weights only, excluding the KV cache.

```python
for name, b in (("7B", 7), ("70B", 70)):
    for label, bytes_per in (("fp16", 2), ("int8", 1), ("4-bit", 0.5)):
        print(name, label, b * bytes_per, "GB")
```

Output:

```
7B fp16 14 GB
7B int8 7 GB
7B 4-bit 3.5 GB
70B fp16 140 GB
70B int8 70 GB
70B 4-bit 35.0 GB
```

**Quiz:** What does 4-bit quantization mainly reduce?

- [ ] The vocabulary
- [ ] The number of users
- [x] Memory used by the weights
- [ ] The training data

*Answer:* Memory used by the weights. Fewer bits per weight means a smaller memory footprint.
