Course 10, lesson 94 of 100, Adults

The transformer, layer by layer

The architecture behind modern LLMs

Like I’m 5

A transformer reads all the words together and lets each word look at the others to understand it better. Then it does this again and again, layer after layer.

The big idea

Text becomes tokens, each token becomes an embedding, and position information is added. A stack of identical blocks follows; each has multi-head self-attention, where every token gathers information from relevant tokens, and a feed-forward network, with residual connections and normalisation around both.

In attention, each token makes a query, key and value. Scores = softmax(QKᵀ/√d), and the output is the weighted sum of values. Decoder models mask future tokens so each position predicts the next one. The final layer outputs a probability for every token in the vocabulary.

Examples

  • Heads: Different heads track grammar, references or topics.
  • Causal mask: When predicting word 5, the model may only look at words 1 to 4.
  • KV cache: Saved keys and values make generation faster, token by token.

How it works

  1. Turn tokens into embeddings and add position information.
  2. Pass them through stacked attention and feed-forward blocks.
  3. Turn the final vectors into next-token probabilities.

Check your understanding

What are the two main parts of each transformer block?
Options: Self-attention and a feed-forward network; A keyboard and a mouse; Convolution and pooling only.
Answer: Self-attention and a feed-forward network. Attention mixes information across tokens; the feed-forward part processes each token.
Why do decoder models use a causal mask?
Options: So each position can't peek at future tokens; To hide the model's name; To save colours.
Answer: So each position can't peek at future tokens. Masking makes next-token prediction honest during training.

Remember

Transformers stack attention and feed-forward blocks to turn tokens into next-token predictions.

Talk about it

Explain to a friend why 'it' in a sentence needs attention to understand.

Go deeper

'Attention Is All You Need' (Vaswani et al., 2017) introduced the transformer. Modern variants use rotary position embeddings, grouped-query attention and mixture-of-experts feed-forward layers.