Course 10, lesson 100 of 100, Adults
Alignment and AI safety
Making AI do what we really intend
Like I’m 5
Alignment means making sure powerful AI tries to do what people really want, safely and honestly, even in situations nobody planned for.
The big idea
Specifying goals is hard: systems optimise what we measure, not what we mean, which leads to reward hacking. As models grow more capable, they must also be robust to tricky inputs, honest about uncertainty and resistant to misuse.
Current techniques include preference tuning, constitutional principles, red-teaming, dangerous-capability evaluations, interpretability and scalable oversight, where AI helps humans supervise AI. Many labs publish safety frameworks that tie stronger safeguards to more capable models.
Examples
- Reward hacking: An agent games its score instead of doing the task.
- Red-teaming: Experts probe for harmful capabilities before release.
- Honesty: Training models to say 'I'm not sure' rather than bluff.
How it works
- Define the intended behaviour and the risks clearly.
- Train with methods that reward helpful, honest, harmless behaviour.
- Test hard with evals and red-teaming, then monitor after release.
Check your understanding
- What is reward hacking?
- Options: Exploiting the measured goal instead of doing what was intended; Stealing rewards from other AIs; Hacking a computer for points.
Answer: Exploiting the measured goal instead of doing what was intended. It shows the gap between what we measure and what we mean. - What is scalable oversight?
- Options: Using AI to help humans supervise more capable AI; Watching AI through a telescope; Letting AI supervise itself with no checks.
Answer: Using AI to help humans supervise more capable AI. It aims to keep human supervision effective as models get smarter.
Remember
Alignment aims to keep capable AI helpful, honest and safe, using training, testing and oversight.
Talk about it
What's one rule you think every powerful AI should follow? Why?
Go deeper
Research areas include specification, robustness, interpretability, oversight and governance. Concrete Problems in AI Safety (Amodei et al., 2016) is a classic starting point.