# Alignment: RLHF and Preference Optimisation — Large Language Models

Source: https://www.geekswithgeeks.com/en/llms/r-align

> Explain how human preferences steer model behaviour.

## Learning from rankings

After SFT, labs collect **preference data**: for one prompt, people (or carefully guided AI judges) say which of two answers is better. In **RLHF** a **reward model** is trained to predict those preferences and the language model is then optimised, with reinforcement learning, to earn higher reward while staying close to the original. Newer methods such as **DPO** learn directly from the preference pairs without a separate reward model. Alignment improves helpfulness and safety but is imperfect: models can still be tricked, can become over-agreeable (**sycophancy**) and can be over-cautious.

**Quiz:** What data does preference tuning use?

- [ ] Only images
- [ ] Only raw web pages
- [x] Comparisons showing which answer is better
- [ ] Server logs

*Answer:* Comparisons showing which answer is better. Preference pairs teach the model which behaviours people favour.
