Lesson 2 / 27

The Life of a Model: Pretraining to Chat

Walk through pretraining, instruction tuning and alignment.

Three broad stages

(1) Pretraining: the model reads trillions of tokens of text and learns to predict the next token; this is by far the most expensive stage and produces a base model that continues text but does not follow instructions well. (2) Supervised fine-tuning (SFT): training on curated examples of instructions and good answers so the model behaves like an assistant. (3) Preference tuning / alignment (RLHF, DPO and related methods): humans or AI judges rank answers and the model is nudged toward helpful, honest and harmless ones. Specific recipes differ between labs and change quickly, but this overall shape is common.

The stages at a glance

Each stage changes what the model is good at.

Stage              Data                         Result
Pretraining        web, books, code (trillions)  base model: fluent, knows a lot, ignores instructions
SFT                instruction/answer pairs      follows instructions, chat format
Preference tuning  ranked answers                more helpful, safer, better tone
Deployment         prompts + tools + retrieval   the product users see

Base versus chat models

Many providers publish both. Use the chat or instruct version for assistants and the base version only if you want raw text continuation or plan to tune it yourself.

Quick check: Which stage is usually the most expensive?

  • Pretraining
  • Writing the prompt
  • Rendering the UI
  • Spell checking
Answer

Pretraining — Pretraining consumes most of the compute because of the data and model size.