Lesson 24 / 27

Model Size and Quantization

Estimate hardware needs and shrink models.

Fewer bits per weight

Model memory is roughly parameters × bytes per parameter. Quantization stores weights with fewer bits (for example 8-bit or 4-bit instead of 16-bit), cutting memory and often speeding inference, with a usually small and task-dependent quality loss. This is what lets a 7B-parameter model run on a laptop or a modest GPU. Other size levers: distillation (train a small model to imitate a big one), mixture-of-experts (only part of the network is active per token), and choosing a smaller model for the task. Remember to leave memory for the KV cache and activations.

Size, speed, choice

Quantization, batching and the open-versus-hosted choice decide cost and control.

Three questions: fit, serve, choose.
Figure 7.1 — Fit, serve and choose.

Weights memory by precision

I ran this plain-Python (standard library only) example. The weights of a 70B model need 140 GB in fp16 but 35 GB at 4-bit; a 7B model drops from 14 GB to 3.5 GB. These are weights only, excluding the KV cache.

for name, b in (("7B", 7), ("70B", 70)):
    for label, bytes_per in (("fp16", 2), ("int8", 1), ("4-bit", 0.5)):
        print(name, label, b * bytes_per, "GB")

Output:

7B fp16 14 GB
7B int8 7 GB
7B 4-bit 3.5 GB
70B fp16 140 GB
70B int8 70 GB
70B 4-bit 35.0 GB

Quick check: What does 4-bit quantization mainly reduce?

  • The vocabulary
  • The number of users
  • Memory used by the weights
  • The training data
Answer

Memory used by the weights — Fewer bits per weight means a smaller memory footprint.