Course 5, lesson 43 of 100, Ages 11+
Samples and fairness
Does your data represent everyone?
Like I’m 5
If you only ask your best friend what everyone's favourite food is, you'll get the wrong answer. A good sample asks lots of different people.
The big idea
Usually you can't collect data from everyone, so you take a sample. A good sample looks like the whole group: different ages, places, languages and backgrounds, in the right proportions.
If a sample leaves people out, the AI works worse for them. Speech AI trained mostly on adult voices may struggle with children. Face recognition trained mostly on some skin tones has made more mistakes on others.
Examples
- Lunch survey: Asking only the football team won't represent the whole school.
- Voice data: Including children, elders and many accents improves speech AI.
- Medical studies: Testing medicines on many groups shows who they work for.
How it works
- Describe the whole group you care about.
- Collect a sample that matches that group's variety.
- Check results separately for each part of the group.
Check your understanding
- What makes a good sample?
- Options: It looks like the whole group in its variety; It's made of your friends only; It's as small as possible.
Answer: It looks like the whole group in its variety. A representative sample reflects everyone you care about. - A speech AI mostly heard adults. Who might it struggle with?
- Options: Children; Adults; Nobody.
Answer: Children. Groups missing from the data often get worse results.
Remember
A good sample represents everyone. Missing groups lead to AI that works worse for them.
Talk about it
If you surveyed your town about parks, who might be easy to leave out?
Go deeper
Sampling bias and selection bias are classic statistical traps. Stratified sampling deliberately includes each subgroup in proportion, and models should be evaluated per subgroup, not just overall.