Lesson 14 / 27
Scaling Laws and Data
Relate parameters, data and compute to model quality.
Bigger helps, but balance matters
Research on scaling laws found that loss falls smoothly and predictably as you increase parameters, training tokens and compute. Studies such as the Chinchilla work showed that, for a given compute budget, parameters and data should grow together; many earlier models were under-trained on too little data. Data quality matters as much as quantity: deduplication, filtering low-quality or toxic text, balancing languages and code, and removing test-set leakage all change results. Training also needs thousands of accelerators and weeks of time, which is why most teams use or adapt existing models instead of training from scratch.
Scale, data, tuning
Capability comes from data and compute; tuning shapes behaviour.
Check for test-set leakage
If benchmark questions appear in the training data, scores look better than real ability. Prefer evaluations you built from fresh data.
Quick check: What did scaling-law research suggest for a fixed compute budget?
- Neither matters
- Only add parameters
- Only add data
- Grow parameters and training data together
Answer
Grow parameters and training data together — Balanced growth of model size and data gives better loss for the same compute.