Course 10, lesson 96 of 100, Adults
RLHF and preference tuning
Teaching models what people prefer
Like I’m 5
After an AI learns to write, people show it which answers they like better. Bit by bit, it learns to give more helpful and kinder answers.
The big idea
Reinforcement learning from human feedback starts with a supervised model. People compare pairs of responses and pick the better one; a reward model learns to predict those preferences. The language model is then optimised to score highly, with a penalty for drifting too far from its original behaviour.
Simpler alternatives like Direct Preference Optimisation train on preference pairs directly. Constitutional AI uses a written set of principles and AI-generated feedback to scale the process. All of these improve helpfulness and safety, but can also teach models to tell people what they want to hear, so evaluation matters.
Examples
- Comparisons: Raters choose between two answers to the same prompt.
- Reward model: A model that predicts which answer people would prefer.
- Sycophancy: Over-optimising for approval can make a model too agreeable.
How it works
- Collect human comparisons between pairs of answers.
- Train a reward model to predict the preferred answer.
- Optimise the language model towards higher reward, staying close to the original.
Check your understanding
- What does the reward model learn in RLHF?
- Options: To predict which answers people prefer; To generate images; To count tokens.
Answer: To predict which answers people prefer. It turns human comparisons into a learnable score. - What's a known risk of optimising for human approval?
- Options: Sycophancy: telling people what they want to hear; Models become too slow; Models forget the alphabet.
Answer: Sycophancy: telling people what they want to hear. Approval isn't always the same as truth, so evals check for it.
Remember
Preference tuning uses human or AI comparisons to steer models towards helpful, safe answers.
Talk about it
If you rated chatbot answers, what would make one answer better than another?
Go deeper
InstructGPT (Ouyang et al., 2022) popularised RLHF with PPO and a KL penalty. DPO (Rafailov et al., 2023) and Constitutional AI (Bai et al., 2022) are widely used alternatives.