Course 10, lesson 97 of 100, Adults

Reinforcement learning

Agents, rewards and policies

Like I’m 5

A robot tries things, gets points for good results, and slowly learns the best way to act. That's reinforcement learning.

The big idea

An agent observes a state, takes an action, receives a reward and moves to a new state. Its goal is a policy, a way of choosing actions, that maximises total future reward, usually with later rewards discounted.

The agent must balance exploration (trying new actions) and exploitation (using what works). Value-based methods like Q-learning estimate how good actions are; policy-gradient methods adjust the policy directly. Deep RL beat world champions at Go and now helps train reasoning models.

Examples

  • Maze: +1 for reaching the exit, −0.01 per step to encourage speed.
  • Exploration: Sometimes trying a new path finds a shortcut.
  • Robotics: Learning to walk in simulation, then transferring to a real robot.

How it works

  1. Observe the current state.
  2. Choose an action using the current policy, sometimes exploring.
  3. Receive a reward, update the policy or values, and repeat.

Check your understanding

What is a policy in reinforcement learning?
Options: A way of choosing actions in each state; A legal document; A type of reward.
Answer: A way of choosing actions in each state. The policy maps states to actions.
Why must an agent explore sometimes?
Options: To discover actions that might be better; To waste time; Exploration is never useful.
Answer: To discover actions that might be better. Without exploring, it may never find better strategies.

Remember

RL agents learn policies that maximise long-term reward by balancing exploration and exploitation.

Talk about it

When do you explore something new instead of sticking with your favourite?

Go deeper

The Bellman equation underlies value methods. Algorithms include DQN, PPO and actor-critic methods. Sim-to-real transfer and reward design remain central challenges.