# Scaling Laws and Data — Large Language Models

Source: https://www.geekswithgeeks.com/en/llms/r-scaling

> Relate parameters, data and compute to model quality.

## Bigger helps, but balance matters

Research on **scaling laws** found that loss falls smoothly and predictably as you increase parameters, training tokens and compute. Studies such as the Chinchilla work showed that, for a given compute budget, parameters and data should grow together; many earlier models were under-trained on too little data. **Data quality** matters as much as quantity: deduplication, filtering low-quality or toxic text, balancing languages and code, and removing test-set leakage all change results. Training also needs thousands of accelerators and weeks of time, which is why most teams use or adapt existing models instead of training from scratch.

## Scale, data, tuning

Capability comes from data and compute; tuning shapes behaviour.

![Three levers: pretrain, fine-tune, align.](assets/figures/llms/section-4-map.svg) — Figure 4.1 — Pretrain, fine-tune and align.

## Check for test-set leakage

If benchmark questions appear in the training data, scores look better than real ability. Prefer evaluations you built from fresh data.

**Quiz:** What did scaling-law research suggest for a fixed compute budget?

- [ ] Neither matters
- [ ] Only add parameters
- [ ] Only add data
- [x] Grow parameters and training data together

*Answer:* Grow parameters and training data together. Balanced growth of model size and data gives better loss for the same compute.
