History of Language Models
See the gap first, then trace where it came from.
Three models, side by side
Run three language models side by side and compare what comes back: a table of counts (1948), a small LSTM (1997), and GPT-2 small (2019), a Transformer with 124M parameters. All three take the same prompt. The n-gram model and the LSTM learn from the same book, Alice's Adventures in Wonderland; GPT-2 read about 8 million web pages (Radford et al., 2019). The difference is obvious before any theory. The rest of the lecture works out where it comes from.
All three designs have been used in products. In 2007 Google's machine translation used n-gram tables with up to 300 billion entries (Brants et al., 2007). In 2016 Google Translate replaced its phrase-based statistical system with an LSTM network (Wu et al., 2016). ChatGPT is fine-tuned from a much larger GPT model (OpenAI, 2022).
Architecture
The Transformer can do something the older models structurally cannot. Testable by building models of matched size and removing pieces.
Scale
Data and compute. Testable with scaling laws: predict the loss before spending the money.
They are not rivals. Both are true, and the useful question is how much each one explains.
History of generative models: from n-grams to the Transformer
- 1948n-gram: Shannon approximates English with letter and word n-grams.
- 1990RNN: the simple recurrent network (Elman).
- 1997LSTM (Hochreiter & Schmidhuber).
- 2003NPLM, the neural probabilistic language model: learned word vectors and an MLP (Bengio et al.).
- 2010RNN language model (Mikolov et al.).
- 2014RNN encoder–decoder (seq2seq, Sutskever et al.), plus attention for translation (Bahdanau et al.): still built on an RNN.
- 2017Transformer: attention replaces the RNN (Vaswani et al.).
- 2018–20GPT, GPT-2, GPT-3: the decoder alone, pre-trained to predict the next token, then scaled up.
Each step below gets the same treatment: the structure, what it fixed, and what it left broken.
n-grams: counting
The Markov assumption keeps only the last \(n-1\) tokens: \[ p(x_t \mid x_1,\dots,x_{t-1}) \approx p(x_t \mid x_{t-n+1},\dots,x_{t-1}) = \frac{\mathrm{count}(x_{t-n+1},\dots,x_{t-1},\textcolor{#ef7c00}{x_t})}{\mathrm{count}(x_{t-n+1},\dots,x_{t-1})} \] There is nothing to learn. The table of counts is the whole model.
NPLM (2003): learn word vectors instead of counting
The neural probabilistic language model of Bengio et al. (2003) gives every word a learned vector, look up the vectors for a fixed window of previous words, and feed them to an MLP. Similar words end up with similar vectors, so what is learned about “cat” carries over to “dog”.
The vector of word \(w\) is \(C(w)\), row \(w\) of a \(|V| \times m\) table \(C\) of learned numbers. For a window of \(n-1\) words the model computes (eq. 1 of the paper, without its optional direct connections)
\[ x = \big(C(w_{t-n+1}), \dots, C(w_{t-1})\big), \qquad h = \tanh(Hx + d), \qquad P(w_t = i \mid \text{context}) = \mathrm{softmax}(Uh + b)_i . \]
\(C, H, d, U, b\) are all trained together by gradient descent. \(H\) is \(h \times (n-1)m\): one block of columns for each position in the window. The window size is built into the weights, so a longer window is a different model, and a word outside the window has no way in.
How neural models learn: gradient descent and backpropagation
An n-gram model is trained by counting. From the NPLM on, every model here has weights that are learned, and training repeats three steps. Compute the loss on some text: the cross-entropy of the true next token from Part 1 (lecture 02), \( \mathcal{L} = -\log p_\theta(x_{t+1} \mid x_{\le t}) \). Compute its gradient \(\nabla_\theta \mathcal{L}\), which says how the loss changes as each weight changes. Move every weight a small step against it, \( \theta \leftarrow \theta - \eta\,\nabla_\theta \mathcal{L} \), where \(\eta\) is the learning rate.
Backpropagation (Rumelhart et al., 1986) computes all the gradients with the chain rule, from the loss back towards the input. For a stack of \(L\) layers with outputs \(h_1, \dots, h_L\), \[ \frac{\partial \mathcal{L}}{\partial h_1} = \frac{\partial \mathcal{L}}{\partial h_L}\,\frac{\partial h_L}{\partial h_{L-1}} \cdots \frac{\partial h_2}{\partial h_1}, \] one factor per layer, or per step in an RNN. If the factors are mostly smaller than 1, the product shrinks towards zero and the early layers barely learn (vanishing gradients); if they are larger than 1, it explodes (Bengio et al., 1994). Several designs in this lecture exist because of this product: the LSTM's cell state and residual connections (C4) give the gradient a path on which the factors stay near 1, and LayerNorm (C5) keeps the size of each layer's input fixed.
RNN: carry a hidden state forward
A recurrent neural network (Elman, 1990) updates one hidden state at every step, \(h_t = \tanh(W h_{t-1} + U x_t)\), so context is no longer limited to a window. The same weights are used at every step.
LSTM: gates that let gradients survive
The LSTM (Hochreiter & Schmidhuber, 1997) adds a cell state that is updated by addition, with gates deciding what to forget, what to write and what to output. When the forget gate (Gers et al., 2000) stays near 1, information and gradients pass through many steps.
Seq2seq + attention (2014–15): a patch
In seq2seq translation (Sutskever et al., 2014) the encoder compresses the whole source sentence into one vector. Attention (Bahdanau et al., 2015) was introduced to let the decoder look back at every encoder state. It was a fix on top of the RNN, not a replacement.
Transformer (2017): keep attention, drop recurrence
The Transformer (Vaswani et al., 2017) is attention and an MLP, stacked, with position added explicitly. GPT-style models use only the decoder half. Click through the diagram; each component has its own section in Part B.
RNN vs Transformer: two numbers explain it
Five architectures, four questions
| Model | Context | Path between two tokens | Parallel training | Why it lost / won |
|---|---|---|---|---|
| n-gram | last \(n-1\) tokens | — | n/a (counting) | No generalisation |
| NPLM (2003) | fixed window | 1 step, inside the window | yes | The window cannot grow |
| RNN | unbounded in principle | \(O(n)\) | no | Forgets; sequential training |
| LSTM | long | \(O(n)\) | no | Memory fixed; bottleneck not |
| Transformer | the full context window | \(O(1)\) | yes | Both at once, at \(O(n^2)\) cost |
Transformers beyond language
The Transformer did not stay in language. Cut the input into tokens and the same Transformer applies; only the two ends change: how the input becomes tokens, and what comes out. One paper per domain:
DiT · Peebles & Xie, 2023
A diffusion model whose denoising network is a Transformer over latent patches, in place of the usual U-Net.
Whisper · Radford et al., 2023
An encoder–decoder, like the 2017 Transformer, from spectrogram to text; trained on 680,000 hours of audio.
PatchTST · Nie et al., 2023
One token per window of the series; forecasts what comes next.
ESM-2 · Lin et al., 2023
One token per amino acid; a language model of up to 15 billion parameters, from which 3D structure is predicted.
The sketches are simplified redrawings, not the papers' figures. Examples: a DiT-XL/2 sample from facebookresearch/DiT (CC BY-NC 4.0), shown in pixels although DiT works on a compressed latent; Eadweard Muybridge, The Horse in Motion (1878, public domain, via Wikimedia Commons); ubiquitin drawn from the coordinates of PDB 1UBQ; the eagle is a U.S. Fish and Wildlife Service photo and the speech is Neil Armstrong’s “That’s one small step for man” (Apollo 11, NASA), both public domain, via Wikimedia Commons.
What makes Transformers great?
One controlled experiment, then the components themselves.
The comparison we just made was not controlled
A Transformer differs from a ConvNet or an RNN in many ways at once: attention, normalisation, activation, optimiser, training recipe, data augmentation, depth ratios. When one model beats another on a dozen axes, the result does not say which axis mattered.
Liu et al. (2022), A ConvNet for the 2020s, ran the controlled version. Start from a ResNet-50 (He et al., 2016), apply the design decisions that came with vision Transformers one at a time, and measure after every change.
Credit for the ConvNet-vs-Transformer gap had gone partly to the wrong place. The training recipe alone was worth +2.7 points, and twelve design changes added up to a large one.
That attention is useless. This is ImageNet classification: fixed-size images, no long-range dependency across a sentence. Language is where one-hop paths pay off.
The six components inside nanoGPT
nanoGPT implements a GPT-2-style model in about 300 lines of model.py. It has six kinds of component, C1–C6. Tokens enter at the bottom, become vectors (C1), pass through n_layer identical blocks (C2–C5), and leave as next-token probabilities (C6). Each code link below opens that part of model.py.
What is it actually trained to do?
Maximum likelihood over the training text, factorised token by token: \[ \theta^* = \arg\max_\theta \sum_i \log p_\theta(x^{(i)}),\qquad \log p_\theta(x) = \sum_t \log p_\theta(x_t \mid x_{<t}) \] Negate and average over tokens to get the loss, and exponentiate to get perplexity: \[ \mathcal{L} = -\frac{1}{T}\sum_t \log p_\theta(x_t \mid x_{<t}),\qquad \mathrm{PPL} = e^{\mathcal{L}} \] Nothing in this lecture changes the objective. The architectures above, and the six components below, are different ways to compute \(p_\theta\).
In training, every position of a sequence predicts its next token at once, and the loss is the average of \(-\log p\) over those predictions: the cross-entropy with the true next token. Gradient descent lowers it (how neural models learn).
C1 · Token and position embeddings
\( x = E[\text{idx}] + P[0,\dots,T-1] \). The token table is the 2003 idea unchanged. The position table is needed because attention has no built-in notion of order.
C2 · Causal self-attention
\[ \mathrm{Attn}(Q,K,V) = \mathrm{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}} + M\right)V,\qquad Q = XW_Q,\; K = XW_K,\; V = XW_V \] This is the only place in the model where tokens exchange information. Remove it and every position is processed independently.
C3 · The MLP, with GELU
\( \mathrm{MLP}(x) = W_2\,\mathrm{GELU}(W_1 x) \) with a hidden width of \(4d\) and the GELU activation (Hendrycks & Gimpel, 2016). It works on each position separately. Most people guess attention holds most of the parameters; the MLP does.
Some nonlinearity is needed, or \(W_2 W_1 x\) collapses to one matrix. Which one is an empirical choice: the 2017 Transformer used ReLU, GPT-1 (Radford et al., 2018) and BERT (Devlin et al., 2019) switched to GELU, and Llama (Touvron et al., 2023) uses SwiGLU (Shazeer, 2020). In ConvNeXt, swapping ReLU for GELU left accuracy unchanged at 80.6% (Liu et al., 2022).
The difference from ReLU is in the slope, which is what a gradient is multiplied by on its way back through the unit. For \(x < 0\), ReLU's slope is exactly 0, so no gradient passes; GELU, \(x\,\Phi(x)\) with \(\Phi\) the standard normal CDF, is smooth, and its slope there is small but not 0 (except at its minimum, near \(x = -0.75\)). Hendrycks & Gimpel reported that it beat ReLU on all the vision, language and speech tasks they tested; GPT-1 adopted it, and BERT notes that it uses GELU “following OpenAI GPT”.
C4 · Residual connections
\( x \leftarrow x + \mathrm{Sublayer}(x) \). Unrolled, \( x_N = x_0 + \sum_n \mathrm{Sublayer}_n(x_n) \): there is a direct path from the input to the output, so in backpropagation the gradient reaches the first layer without being multiplied by one factor per layer (see above). Residual connections came from ResNet (He et al., 2016) and are still in every large model.
C5 · LayerNorm, and where to put it
LayerNorm (Ba et al., 2016): \( \mathrm{LN}(x) = \gamma \odot \dfrac{x-\mu}{\sqrt{\sigma^2+\epsilon}} + \beta \), where \(\mu\) and \(\sigma^2\) are the mean and variance of the \(d\) numbers in one token's vector, and \(\gamma\), \(\beta\) are learned. The 2017 Transformer (Vaswani et al., 2017) normalised after each residual addition; GPT-2 (Radford et al., 2019) and nanoGPT normalise before each sub-layer, which leaves the residual path untouched and makes deep stacks easier to train.
C6 · The output head and weight tying
\( \text{logits} = \mathrm{LN}(x)\,E^\top \). nanoGPT uses the embedding matrix again as the output layer (weight tying, Press & Wolf, 2017): the same space in, the same space out, and \(V\times d\) fewer parameters.
What nanoGPT does not have — yet
| Component | What it changes |
|---|---|
| RoPE | Rotary position embedding: relative position, no learned table |
| RMSNorm | LayerNorm without subtracting the mean; cheaper, about as good |
| SwiGLU | A gated MLP; the hidden width shrinks to \(8/3\cdot d\) to keep the parameter count |
| Grouped-query attention | Fewer key/value heads than query heads; a smaller memory cache at inference |
| Mixture of experts | Many MLPs, only a few active for each token |
All five are refinements of C1–C6 found in current open models. Session 02 returns to them.
Break a GPT
The ConvNeXt method, on a model small enough to run on a laptop.
Ablation, live
What we break. nanoGPT is a minimal GPT: a model file of about 300 lines with the six components above. We use a one-file copy with the same six parts, made small on purpose: TinyShakespeare (char-rnn), 1.1 MB of text with one character per token, and a model of 818,241 parameters, about 200,000 times fewer than GPT-3 (Brown et al., 2020). One training run takes about a minute on a laptop. Unlike nanoGPT, the copy does not tie the output head to the token embedding and has no dropout.
Data: TinyShakespeare, one character per token (65 characters); first 90% to train, last 10% to validate. Model: 4 blocks, 4 heads, width 128, context 64 characters, 818,241 parameters. Training: 3,000 steps of 32 sequences; AdamW, learning rate 3 × 10⁻⁴; gradients clipped at 1.0. Seeds: 1337, 1338, 1339. Measure: validation loss, averaged over 20 batches.
Remove LayerNorm · remove residual connections · remove positional embeddings · collapse multi-head attention to one head.
Reading an ablation honestly
- Train each version more than once. Use a different random seed each time. One run can be lucky or unlucky, so report the average and the range.
- Train every version the same way. Same number of steps, same data. Otherwise a difference may come from the training, not from the change.
- Say what else the change removed. Taking out a part can also take out parameters, so the model is smaller as well as different.
- Report changes that made no difference. “It made no measurable difference” is a real finding.
Homework, for practice
Run your own ablation on the small GPT. The goal: find out what one component does, by training the model with it and without it.
- Run the baseline: the base code below as given, 3 seeds.
- Change one component, nothing else: remove it or replace it, with the same data, steps and seeds. Choose one of today's four in more depth (does removing LayerNorm still help at 12 layers?), another (MLP width, depth, GELU → ReLU, tied vs untied embeddings), or add one nanoGPT lacks (RMSNorm, RoPE).
- Compare and explain: is the change in validation loss larger than the spread across seeds? What does the component do that explains it?
What you produce:
- Your code change: a diff against
ablation_base.py, with its fixed settings unchanged. - A table: validation loss, mean and range over 3 seeds, for the baseline and your version.
- One plot: the training-loss curves of both.
- A few sentences: what the component does, why the loss changed or did not, and whether you predicted it.
Before you conclude, check your runs against the four rules above, and check that the loss at step 0 is near ln 65 = 4.17.
Base code: ablation_base.py, the model and settings behind the class table: a nanoGPT-shaped Transformer (Karpathy, 2022) in one file, with the fixed settings at the top and the four class variants as switches. Add your variant as a new switch, then run python ablation_base.py --variant baseline and --variant your_variant; each runs the three seeds. The output file also holds the training loss every 10 steps, for your plot.
python -m venv gpt && source gpt/bin/activate # optional: a clean environment
pip install torch # Python 3.9+; no GPU needed
curl -O https://costa-nus.github.io/CEG5305_GenAI/part2/lecture-1/code/ablation_base.py
python ablation_base.py --out baseline.json # 3 seeds; downloads the data
# your change: add one entry to VARIANTS in ablation_base.py, e.g. "deeper": {"n_layer": 12}
python ablation_base.py --variant deeper --out mine.json
Where to run it: open the notebook in Colab (choose a T4 GPU runtime; about 35 s per run), or run it locally with Python 3.9+ and PyTorch. No GPU is needed: the run uses under 0.5 GB of RAM and takes about 75 s per run on a laptop CPU (measured on an Apple M5 Pro), so about 20 minutes for all 15 runs of the class table.
Resources
Reading
- Vaswani et al., Attention Is All You Need, NeurIPS 2017. §3, Model Architecture; read §3.5 on positional encoding carefully. Read it as a historical document: most current models differ from it in several places this page points out.
- Liu et al., A ConvNet for the 2020s, CVPR 2022. §2, the modernisation roadmap.
- Raschka, Build a Large Language Model (From Scratch), Manning 2024. Chapters 2–4. Code: rasbt/LLMs-from-scratch.
The official reading list for the module is on Canvas.
Code
- karpathy/nanoGPT — the model this lecture takes apart (Karpathy, 2022).
- karpathy/minGPT — the same model written for reading rather than training.
- karpathy/minbpe — byte-pair encoding tokenisation.
More demos
- AttentionViz — query and key vectors of trained models across layers and heads. Heavy: takes up to a minute to load.
References
Every work cited on this page, in APA style with shortened author lists. Titles link to the paper.
Next session: how the model is actually trained
The objective at scale
What it takes to minimise the next-token prediction loss over trillions of tokens.
Pre-training data
Where the tokens come from, and why most of them are thrown away.
Scaling laws
Predicting that loss from model size, data and compute, before spending the compute.
Adapting without retraining
LoRA and parameter-efficient fine-tuning.