Lesson 16 / 27

Alignment: RLHF and Preference Optimisation

Explain how human preferences steer model behaviour.

Learning from rankings

After SFT, labs collect preference data: for one prompt, people (or carefully guided AI judges) say which of two answers is better. In RLHF a reward model is trained to predict those preferences and the language model is then optimised, with reinforcement learning, to earn higher reward while staying close to the original. Newer methods such as DPO learn directly from the preference pairs without a separate reward model. Alignment improves helpfulness and safety but is imperfect: models can still be tricked, can become over-agreeable (sycophancy) and can be over-cautious.

Quick check: What data does preference tuning use?

  • Only images
  • Only raw web pages
  • Comparisons showing which answer is better
  • Server logs
Answer

Comparisons showing which answer is better — Preference pairs teach the model which behaviours people favour.