Lesson 14 / 27

Scaling Laws and Data

Relate parameters, data and compute to model quality.

Bigger helps, but balance matters

Research on scaling laws found that loss falls smoothly and predictably as you increase parameters, training tokens and compute. Studies such as the Chinchilla work showed that, for a given compute budget, parameters and data should grow together; many earlier models were under-trained on too little data. Data quality matters as much as quantity: deduplication, filtering low-quality or toxic text, balancing languages and code, and removing test-set leakage all change results. Training also needs thousands of accelerators and weeks of time, which is why most teams use or adapt existing models instead of training from scratch.

Scale, data, tuning

Capability comes from data and compute; tuning shapes behaviour.

Three levers: pretrain, fine-tune, align.
Figure 4.1 — Pretrain, fine-tune and align.

Check for test-set leakage

If benchmark questions appear in the training data, scores look better than real ability. Prefer evaluations you built from fresh data.

Quick check: What did scaling-law research suggest for a fixed compute budget?

  • Neither matters
  • Only add parameters
  • Only add data
  • Grow parameters and training data together
Answer

Grow parameters and training data together — Balanced growth of model size and data gives better loss for the same compute.