Course 4, lesson 40 of 100, Ages 10+

How AI recognises images

Filters, layers and convolutions

Like I’m 5

Imagine sliding a small window over a picture, looking for a tiny pattern like an edge. AI does this with lots of little windows, then puts the clues together.

The big idea

Image AI often uses convolutional neural networks. A small filter, maybe 3 by 3 pixels, slides across the image and lights up wherever its pattern appears, such as a vertical edge or a curve.

Layer after layer, these filters combine into detectors for textures, then parts like wheels or ears, then whole objects. Because the same filter is used everywhere, the network can spot a cat whether it's in the corner or the middle.

Examples

  • Edge filter: Lights up where dark meets light, like the outline of a door.
  • Medical scans: Filters learn patterns that could signal disease.
  • Self-driving: Networks find lanes, signs and people in camera images.

How it works

  1. Small filters slide across the image to find simple patterns.
  2. Deeper layers combine those patterns into parts.
  3. The final layers decide what the whole object is.

Check your understanding

What does a filter in an image network do?
Options: Slides across the image looking for a small pattern; Changes the photo to black and white; Deletes the background.
Answer: Slides across the image looking for a small pattern. Each filter detects a pattern wherever it appears.
Why can a network spot a cat anywhere in the picture?
Options: The same filters are used across the whole image; It only looks in the middle; Cats are always in the corner.
Answer: The same filters are used across the whole image. Sharing filters across positions makes detection work everywhere.

Remember

Image AI slides small filters over pictures and builds simple patterns into whole objects.

Talk about it

Look around the room. What simple shapes make up the objects you see?

Go deeper

Convolutional neural networks (CNNs) use weight sharing and pooling for translation-tolerance. Vision transformers instead split images into patches and use attention, and now rival CNNs on many tasks.