Course 10, lesson 100 of 100, Adults

Alignment and AI safety

Making AI do what we really intend

Like I’m 5

Alignment means making sure powerful AI tries to do what people really want, safely and honestly, even in situations nobody planned for.

The big idea

Specifying goals is hard: systems optimise what we measure, not what we mean, which leads to reward hacking. As models grow more capable, they must also be robust to tricky inputs, honest about uncertainty and resistant to misuse.

Current techniques include preference tuning, constitutional principles, red-teaming, dangerous-capability evaluations, interpretability and scalable oversight, where AI helps humans supervise AI. Many labs publish safety frameworks that tie stronger safeguards to more capable models.

Examples

  • Reward hacking: An agent games its score instead of doing the task.
  • Red-teaming: Experts probe for harmful capabilities before release.
  • Honesty: Training models to say 'I'm not sure' rather than bluff.

How it works

  1. Define the intended behaviour and the risks clearly.
  2. Train with methods that reward helpful, honest, harmless behaviour.
  3. Test hard with evals and red-teaming, then monitor after release.

Check your understanding

What is reward hacking?
Options: Exploiting the measured goal instead of doing what was intended; Stealing rewards from other AIs; Hacking a computer for points.
Answer: Exploiting the measured goal instead of doing what was intended. It shows the gap between what we measure and what we mean.
What is scalable oversight?
Options: Using AI to help humans supervise more capable AI; Watching AI through a telescope; Letting AI supervise itself with no checks.
Answer: Using AI to help humans supervise more capable AI. It aims to keep human supervision effective as models get smarter.

Remember

Alignment aims to keep capable AI helpful, honest and safe, using training, testing and oversight.

Talk about it

What's one rule you think every powerful AI should follow? Why?

Go deeper

Research areas include specification, robustness, interpretability, oversight and governance. Concrete Problems in AI Safety (Amodei et al., 2016) is a classic starting point.