Lesson 24 / 27
Model Size and Quantization
Estimate hardware needs and shrink models.
Fewer bits per weight
Model memory is roughly parameters × bytes per parameter. Quantization stores weights with fewer bits (for example 8-bit or 4-bit instead of 16-bit), cutting memory and often speeding inference, with a usually small and task-dependent quality loss. This is what lets a 7B-parameter model run on a laptop or a modest GPU. Other size levers: distillation (train a small model to imitate a big one), mixture-of-experts (only part of the network is active per token), and choosing a smaller model for the task. Remember to leave memory for the KV cache and activations.
Size, speed, choice
Quantization, batching and the open-versus-hosted choice decide cost and control.
Weights memory by precision
I ran this plain-Python (standard library only) example. The weights of a 70B model need 140 GB in fp16 but 35 GB at 4-bit; a 7B model drops from 14 GB to 3.5 GB. These are weights only, excluding the KV cache.
for name, b in (("7B", 7), ("70B", 70)):
for label, bytes_per in (("fp16", 2), ("int8", 1), ("4-bit", 0.5)):
print(name, label, b * bytes_per, "GB")
Output:
7B fp16 14 GB 7B int8 7 GB 7B 4-bit 3.5 GB 70B fp16 140 GB 70B int8 70 GB 70B 4-bit 35.0 GB
Quick check: What does 4-bit quantization mainly reduce?
- The vocabulary
- The number of users
- Memory used by the weights
- The training data
Answer
Memory used by the weights — Fewer bits per weight means a smaller memory footprint.