Part 2 - Lecture 1

Architecture of Large
Language Models

Session 01 of 2 · Architecture (session 02: Training)

CEG5305 Introduction to Generative AI with Foundation Models
Junyuan Hong · AY2026/27 Semester 1

2026.10.01v1

Recap · Part 1, with Liu Xingyu

Where we are: what Part 1 covered

LectureTopicKey message
01Generative and autoregressive modelsA generative model learns p(x); the chain rule turns it into next-token predictions, and a Transformer with a causal mask computes them
02Learning by maximum likelihoodTraining minimises the KL divergence to the data, which means minimising the cross-entropy of each true next token
03Latent variable models, the VAEExplain data with hidden variables z; the ELBO makes training possible
04Diffusion modelsAdd noise step by step, then learn to remove it: DDPM, and DiT and Stable Diffusion in practice
05Normalizing flows, flow matchingInvertible maps give an exact likelihood; flow matching learns the path from noise to data directly

Today builds on 01 and 02: the chain rule, the causal mask and the cross-entropy loss all come back.

Looking ahead · Part 2, with Junyuan Hong

Where we are going: Part 2

LectureTopicKey message
1 · 01Architecture of LLMs todayWhat a GPT is made of, and how to find out which parts matter
1 · 02Training, scaling laws, LoRAA fixed compute budget: how much data, how big a model; then adapting a trained model cheaply
2 · 01Post-training: instruction tuning, RLHFA base model continues text; post-training teaches it to answer
2 · 02Prompting, in-context learning, chain of thoughtChanging what a model does without changing a single weight, and where that stops working
3Self-supervised foundation modelsLearning from unlabelled data by predicting one part from another, beyond text
4Large multi-modal modelsConnecting a vision encoder to an LLM so that it can see

From what the model is, to how it is trained, aligned, prompted, and extended beyond text.

Outline · Lecture 1 = Topic 4, six hours

Two sessions: architecture, then training

Session 01 · today · 3 h

Architecture — what an LLM is made of

☕Break10 min
☕Break10 min
Session 02 · next · 3 h

Training — how the model is actually trained

Training on large datasets · pre-training data · scaling laws · LoRA fine-tuning

Today's question

Why do LLMs work so well?

And the part of that question we can test today:

Is it the architecture, or the scale?

Today we test model design. Session 02 examines model size, data and compute.

Today's question · the record so far

Larger models improve performance. Better designs reduce compute.

Still open: each new design arrived with more compute and data. For 2012–2023 language models, Ho et al. (2024) credit 60–95% of the gains to compute, 5–40% to algorithms.
Why it matters: using 1/7 or 1/11 of the compute means less computation for the same score.

Scores from the papers' tables (Vaswani T2, Devlin T1, Touvron 2023a T9 and 2023b T19, DeepSeek-AI T3, OpenAI T2); compute of the MMLU models from Epoch AI, mostly estimates. *GPT-4: estimated compute (bar: 90% interval), post-trained score. Logos: Simple Icons.

Part A

History of Language Models

Compare the models’ outputs, then see how their designs changed.

1 · Show

Arch makes generation different

1 · Counting

An n-gram model (1948)

Counts from Alice in Wonderland.

Used in: Google's translator in 2007.

2 · Recurrent

An LSTM (1997)

Learned from the same book.

Used in: Google Translate from 2016.

3 · Transformer

GPT-2 small (2019)

124M parameters, web text.

Used in: larger GPT models run ChatGPT.

Example prompt

“The Eiffel Tower is”

✋ Your turn

Suggest a prompt for all three.

1 · Show · demo 1

An n-gram model (1948 · 78 years ago)

Word or character n-grams, with the top-5 next-token probabilities.

1 · Show · demo 2

An LSTM (1997 · 29 years ago)

Reads and writes one character at a time. Trained on Alice in Wonderland only.

1 · Show · demo 3

A Transformer: GPT-2 small (2019 · 7 years ago)

Runs on this laptop. Nothing is sent to a server.

1 · Show

What actually differs?

n-gram · counts from Alice
the eiffel tower is such a large letters it mouse to think about this morning i could the garden and beg your pardon cried alice not
LSTM · learned from Alice
The Eiffel Tower isself cats court of the side, who was going to the arch-arching of tell, which handed the trembled on ord her. "That's are the beginnous way and she co
GPT-2 small · web text
The Eiffel Tower is now an international landmark on the Eiffel Tower estate, where it has been restored to its historic glory. It was originally built in 1923 by the British builder, Sir Winston Churchill. …
✋ What are the differences? Style?

Pre-run on 29 Sep 2026 with the prompt “The Eiffel Tower is”, the first run of each, not chosen for quality. n-gram: word bigram, 22 words. LSTM: temperature 0.7, 150 characters. GPT-2 small: temperature 0.8, top-k 40, the first two sentences of 50 tokens. All three from the demos on the previous slides.

1 · Show

Two hypotheses

Hypothesis 1 · today

Architecture

The Transformer can directly connect distant tokens and process positions in parallel during training.

Hypothesis 2 · session 02

Scale

More training data and more computation.

Both contribute. How much does each explain the differences we observed?

Segment 2 · 25 min

History of generative models
from n-grams to the Transformer

n-gram → NPLM → RNN → LSTM → Transformer

2 · History of generative models · the common principle

The task of LM is to generate a sequence of tokens

Our goal is to generate a sentence \([x_1, \dots, x_T]\), one token at a time.

Chain rule: exact for any sequence.

\[ p(x_1,\dots,x_T) = \prod_{t=1}^{T} p(x_t \mid x_1,\dots,x_{t-1}) \]

For text, code, DNA or music, these next-token probabilities determine the probability of a complete sequence. Generate a sequence by sampling one token at a time.

Markov chain: keep only the last \(k\) tokens.

\[ p(x_t \mid x_1,\dots,x_{t-1}) \approx p(x_t \mid x_{t-k},\dots,x_{t-1}) \]

With \(k\) as long as the sequence, it is exact again, but there are \(V^k\) contexts to learn.

Markov (1913) counted vowel–consonant pairs in Pushkin's Eugene Onegin; Shannon (1948) used such chains to generate English.

2 · History of generative models

Seventy years in eight steps

  1. 1948n-gram: Shannon approximates English by counting
  2. 1990RNN: the simple recurrent network (Elman)
  3. 1997LSTM: gates on a cell state (Hochreiter & Schmidhuber)
  4. 2003NPLM: word vectors + an MLP (Bengio et al.)
  5. 2010RNN language model (Mikolov et al.)
  6. 2014RNN encoder–decoder (Sutskever et al.), plus attention (Bahdanau et al.): still an RNN
  7. 2017Transformer: attention replaces the RNN (Vaswani et al.)
  8. 2018–20GPT, GPT-2, GPT-3: the Transformer’s decoder only, next-token prediction, scaled up

2 · History of generative models

n-gram models predict from counts

Keep only the last \(n-1\) tokens, then count:

\[ p(x_t \mid x_{1},\dots,x_{t-1}) \approx \frac{\mathrm{count}(x_{t-n+1},\dots,x_{t-1},\textcolor{#ef7c00}{x_t})}{\mathrm{count}(x_{t-n+1},\dots,x_{t-1})} \]

The model stores counts from the training text rather than learned neural-network weights.

WORKEDFast, with predictions explained by counts. Used in speech and translation for decades: Google’s 2007 translation system counted 5-grams in 2 trillion tokens (Brants et al., 2007).
BROKECannot share patterns between similar contexts. Most contexts never appear in the data.

2 · History of generative models

An n-gram model is a table of counts

Counted live from the seven sentences shown. Try the trigram model, then click “dog”, “mat”.

Board · derivation 1 · ~5 min

Why n-gram tables cannot scale

✋ “the cat sat on the ___” and “the dog sat on the ___”: does the table know these are related?

2 · History of generative models

NPLM (2003): learn word vectors instead of counting

Word vector: \(C\) is a \(|V| \times m\) table of learned numbers; the vector of word \(w\) is its row, \(C(w)\).

\( x = \big(C(w_{t-n+1}), \dots, C(w_{t-1})\big) \qquad h = \tanh(Hx + d) \qquad P(w_t = i \mid \text{context}) = \mathrm{softmax}(Uh + b)_i \)

Eq. 1 of Bengio et al. (2003), with its optional direct connections \(Wx\) left out. \(C, H, d, U, b\) are trained together by gradient descent. Random, untrained weights here.

2 · History of generative models

What NPLM fixed

A feed-forward network (MLP) reads a fixed window of learned word vectors. Similar words get similar vectors.

FIXEDSimilar word vectors let patterns learned for “cat” also help predict after “dog”.
DID NOTThe window is fixed: H has weights for each position, and earlier words are ignored.

2 · History of generative models · how neural models learn

Training: gradient descent and backpropagation

From the NPLM on, weights are trained, not counted: move each weight against its gradient.

loss \( \mathcal{L} = -\log p_\theta(x_{t+1} \mid x_{\le t}) \) update \( \theta \leftarrow \theta - \eta\,\nabla_\theta \mathcal{L} \) backprop \( \dfrac{\partial \mathcal{L}}{\partial h_1} = \dfrac{\partial \mathcal{L}}{\partial h_L}\,\dfrac{\partial h_L}{\partial h_{L-1}} \cdots \dfrac{\partial h_2}{\partial h_1} \)

Loss: Part 1, lecture 02. Backpropagation: Rumelhart et al. (1986). Left: \(L(w) = \tfrac12 (w-1)^2\). Right: one number \(g\) stands for the size of each factor; an illustration, not a trained model. Vanishing gradients in RNNs: Bengio et al. (1994).

2 · History of generative models

RNN: update a memory vector at each step

\( h_t = \tanh(W\,h_{t-1} + U\,x_t), \qquad y_t = \mathrm{softmax}(V\,h_t) \) \(x_t\): the word vector of token \(t\). \(h_t\): the memory, the same size at every step.

NPLM (2003): a fixed window of word vectors
RNN: one word vector at a time, into the memory \(h\)

Left: the NPLM sees only its window. Right: a 6-unit tanh RNN with random weights; the bars show ‖∂hₜ/∂x₁‖, how much the first input can still change the memory.

2 · History of generative models

What the RNN fixed

\( h_t = \tanh(W h_{t-1} + U x_t) \): one state, the same weights at every step.

FIXEDThe memory vector can, in principle, retain information from any earlier input.
DID NOTGradients from early inputs shrink or grow too large. Training still processes steps in order.

2 · History of generative models

LSTM: gates control what memory keeps

\( c_t = f_t \odot c_{t-1} + i_t \odot g_t, \qquad h_t = o_t \odot \tanh(c_t) \) gates \(f_t, i_t, o_t = \sigma(\cdot)\), between 0 and 1; new content \(g_t = \tanh(\cdot)\); each computed from \(h_{t-1}\) and \(x_t\).

RNN: \(h_t = \tanh(W h_{t-1} + U x_t)\)
LSTM: the cell state \(c\) is changed only by \(\times f\) and \(+\)

After Graves (2013), Fig. 2. Scalar version of one cell.

2 · History of generative models

LSTM: what is left of the cell state after T steps

Forget gate f held fixed, nothing added. Dashed: an RNN shrinking by 0.6 per step (stylised).

2 · History of generative models · demo

An LSTM, writing by hand

What it generates: the next pen move. At each step the LSTM outputs a probability distribution over how far the pen moves in x and y, and whether it lifts off the page. It samples one move, draws it, and reads it back in.

One move at a time (illustration)
The real model: type a phrase, then press Write.

Graves (2013), §4: pen offsets from a mixture of 2-D Gaussians plus an end-of-stroke probability; §5: a soft window over the text tells it which letter it is writing. Left: the pen path is the Hershey script font, not a model sample; the orange ellipses show the kind of distribution output at each step.

2 · History of generative models

What the LSTM fixed, and what it did not

FIXEDMemory over hundreds of steps: the cell state is updated by addition, so while the forget gate stays near 1, information and gradients pass through.
Still a bottleneckConsequence
① Step \(t\) waits for step \(t-1\): \(n\) steps in a row, in training too. Slow training. The \(n\) positions cannot be computed at the same time, so training time grows with sequence length however many GPU cores are free. It is not more arithmetic: per layer, recurrence costs \(O(n \cdot d^2)\) and attention \(O(n^2 \cdot d)\). The difference is \(O(n)\) sequential steps against \(O(1)\).
② Everything about the past must fit in one fixed-size state. Limited capacity. A long input is squeezed into the same \(d\) numbers, so detail is lost, and a word \(n\) steps back reaches the output only through \(n\) updates. A plain encoder–decoder translated worse as sentences got longer.

Costs per layer and sequential steps: Vaswani et al. (2017), Table 1. Longer sentences, lower BLEU: Bahdanau et al. (2015), Fig. 2.

2 · History of generative models

Seq2seq (2014): a whole sentence in one vector

Bottleneck ②: the encoder’s last state \(c\) is all the decoder gets (Sutskever et al., 2014). Attention (Bahdanau et al., 2015) lets each decoder step look back at every encoder state.

\( \alpha_{ij} = \mathrm{softmax}_j\big(a(s_{i-1}, h_j)\big), \qquad c_i = \sum_j \alpha_{ij}\, h_j, \qquad s_i = f(s_{i-1}, y_{i-1}, \textcolor{#ef7c00}{c_i}) \) Score each \(h_j\) against \(s_{i-1}\), turn the scores into weights \(\alpha_{ij}\), feed their weighted average \(c_i\) to step \(i\).

Formulas: Bahdanau et al. (2015), §3.1, Eqs. 4–6; \(a\) is a small learned network. Alignment weights in the diagram are hand-set for illustration, not from a trained model.

2 · History of generative models

RNN vs attention: recurrent update vs a direct connection

Simplified RNN: each update retains a fraction g, as in the LSTM example. Attention weights are set by hand.

2 · History of generative models

Transformer (2017): attention is all you need

(1) seq2seq → encoder–decoder   (2) recurrence → self-attention

Encoder: reads the whole source at once; every word attends to every other. Word order comes from a position vector added to each word, not from reading left to right.

Decoder: writes the target. Cross-attention reads the encoder, the seq2seq attention without the RNN. Its own self-attention is masked: why, on the next slide.

After Fig. 1 and §3.1 of Vaswani et al. (2017), simplified.

2 · History of generative models

Self-attention in one matrix product, and the causal mask

\( \mathrm{softmax}(QK^\top/\sqrt{d} + M)\,V \), with \(M_{ij} = -\infty\) for \(j > i\), else 0: Vaswani et al. (2017), §3.2.3. Weights hand-set; Part B computes it with real numbers.

Board · derivation 2 · ~6 min

RNN vs attention: path length and training steps

One at a time, or all at once: an RNN is the first robot, attention the second. Video: YouTube, LEARN & FUN.

Attention compares O(n²) token pairs, but computes those comparisons in parallel. An RNN must process n steps in order.

2 · History of generative models

Decoder-only (GPT, 2018): drop the encoder, keep the mask

Encoder–decoder · 2017

Trains on paired data

From a French sentence, predict its English translation.

Decoder-only · GPT

Trains on any text

Predict the next token, at every position.

  • Simple architecture: one module, the decoder.
  • A simple causal mask gives the order that recurrence gave the RNN: position \(i\) sees only positions \(\le i\).

GPT-1: a “12-layer decoder-only transformer with masked self-attention heads” (Radford et al., 2018). Translation by prompt, “english sentence = french sentence”: GPT-2 5 BLEU on WMT-14 En→Fr (Radford et al., 2019, §3.7); GPT-3 few-shot 32.6 (Brown et al., 2020, Table 3.4).

2 · History of generative models

The complete Transformer: N identical blocks

Decoder-only: nanoGPT model.py. Encoder–decoder: after Fig. 1 of Vaswani et al. (2017).

2 · History of generative models

What the Transformer fixed

Each block combines attention and a feed-forward network (MLP). Position vectors encode token order.

FIXEDDirect connections between tokens, with all positions processed in parallel during training.
DID NOTAttention compares every pair: time and memory grow as \(n^2\).

2 · History of generative models · demo

A real GPT-2, animated

Type a prompt, press Generate.

2 · History of generative models · demo

nano-gpt, GPT-2 and GPT-3 in 3D

nano-gpt, GPT-2 and GPT-3 side by side. Use full screen.

2 · History of generative models

Five architectures, four questions

ModelContextPath lengthParallel trainingMain strength or limit
n-gramn−1 tokens—n/aCannot share similar contexts
NPLMfixed window1, in windowyesWindow cannot grow
RNNunbounded*O(n)noForgets; sequential
LSTMlongO(n)noLonger memory; still sequential
Transformerfull windowO(1)yesDirect access; O(n²) comparisons

* In principle. Shrinking gradients make learning from distant tokens difficult.

2 · History of generative models

Transformers beyond language

Transformers also process non-text tokens. Each task uses its own input encoding and output layer.

Image generation

DiT · Peebles & Xie (2023)

Video

ViViT · Arnab et al. (2021)

Speech

Whisper · Radford et al. (2023)

Time series

PatchTST · Nie et al. (2023)

Proteins

ESM-2 · Lin et al. (2023)

Examples: eagle, USFWS; Armstrong, Apollo 11, NASA; Muybridge (1878); all public domain. DiT sample, facebookresearch/DiT (CC BY-NC 4.0). Ubiquitin, PDB 1UBQ.

Part A — what attention changes

Each token reads earlier tokens directly. Training processes all positions at once.

These explain advantages over RNNs. Part B tests how other design choices improve performance.

BREAK 1
10 minutesBack at 19:00

Coming up in Part B: an experiment that changes one thing at a time, then nanoGPT’s six components.

Part B

What makes Transformers great?

Test one design change at a time, then examine nanoGPT’s components.

3 · The ConvNeXt experiment

The comparison we just made was not controlled

A Transformer differs from a ConvNet in many ways at once:

  1. 1self-attention instead of convolution
  2. 2LayerNorm (LN) instead of BatchNorm (BN)
  3. 3GELU instead of ReLU
  4. 4AdamW, longer training, more varied images (not drawn)
  5. 5an input layer that cuts the image into patches
  6. 6different numbers of blocks in each stage

If two models differ in many ways, a higher score does not tell us which change helped.

3 · The ConvNeXt experiment

How to investigate which component is important?

1

Start with ResNet-50: 76.1% ImageNet accuracy.

2

Add one design or training change at a time.

3

Measure accuracy after each change, using similar amounts of computation.

Liu et al. (2022), A ConvNet for the 2020s. Diagram schematic: effect sizes made up; the measured ones are on the next slide.

3 · The ConvNeXt experiment

ConvNeXt: one change at a time, measured

Liu et al. (2022), Fig. 2 and §2. ImageNet-1K top-1, ResNet-50 → ConvNeXt-T; Swin-T as the target.

3 · The ConvNeXt experiment

What ConvNeXt does, and does not, show

It shows (click one)

Some gains came from changes besides attention.

Training changes alone: +2.7 percentage points in accuracy.

Twelve design changes together: +3.2 more.

It does not show

That attention is useless.

Whether the same gains apply to language: this experiment tests image classification.

We can use the same experimental method. Change one thing, hold the rest fixed, measure. Part C does it to a GPT.

Segment 4 · 28 min

The components, in nanoGPT

From the paper to about 300 lines of model.py

4 · The components

nanoGPT in one diagram: six components

Tokens become vectors (C1), pass through n_layer blocks with the same structure (C2–C5), then become next-token probabilities (C6). First, what it is trained to do; then each component.

Board · derivation 3 · ~7 min

What is it actually trained to do?

  1. Start from maximum likelihood: \( \theta^* = \arg\max_\theta \sum_i \log p_\theta(x^{(i)}) \)
  2. Expand the sequence probability into next-token probabilities. Taking logs turns their product into a sum.
  3. \( \mathcal{L} = -\frac{1}{T}\sum_t \log p_\theta(x_t \mid x_{<t}) \): token-level cross-entropy.
  4. \( \mathrm{PPL} = e^{\mathcal{L}} \). ✋ What does that number mean?

All these models learn to predict the next token. Their designs, the six components next, change how they compute \(p_\theta\).

4 · The components

The loss: predict every next token

Every position predicts the token after it; the loss averages −log p of the true next token. Training lowers it by gradient descent.

At position t\( H_t = -\sum_{v \in V} y_{t,v} \log p_\theta(v \mid x_{\le t}) = -\log p_\theta(x_{t+1} \mid x_{\le t}) \) \(y_t\): 1 for the true next token, 0 for all others. Over the sequence\( \mathcal{L} = \frac{1}{T}\sum_{t=1}^{T} H_t = -\frac{1}{T}\sum_{t=1}^{T} \log p_\theta(x_{t+1} \mid x_{\le t}) \) the average cross-entropy over positions.

4 · The components

C1 · Token and position embeddings

The same word gets a different vector at a different position, because a position vector is added.

\( x_t = E[\mathrm{idx}_t] + P[t] \) \(E\): \(V \times d\) token table. \(P\): \(T \times d\) position table.

4 · The components

C2 · Causal self-attention: queries, keys, scores

A score compares one token’s query vector with another’s key vector. Higher scores give that token more attention.

\( Q = XW_Q,\; K = XW_K,\; V = XW_V, \qquad S = QK^\top/\sqrt{d_k} \) \(X\): one row per token. \(S_{ij}\): query \(i\) against key \(j\).

Five tokens, d = 4, random weights.

4 · The components

C2 · Mask, softmax, weighted sum

Each output averages its own and earlier tokens’ value vectors, using attention weights. This is where information passes between tokens.

\( A = \mathrm{softmax}(S + M), \qquad \mathrm{out} = A\,V \) \(M_{ij} = -\infty\) for \(j > i\), else 0. Each row of \(A\) sums to 1.

4 · The components

C3 · The MLP, with GELU

Attention mixes tokens; the MLP then updates each token on its own. Its nonlinearity, GELU, is a smooth ReLU whose slope is not zero for negative inputs, so gradients still pass.

\( \mathrm{MLP}(x) = W_2\,\mathrm{GELU}(W_1 x) \) \(W_1\): \(4d \times d\), \(W_2\): \(d \times 4d\). Where are most of the parameters?

4 · The components

C4 · Residual connections

Each sub-layer adds its result to x, so gradients flow straight down x instead of shrinking layer by layer.

\( x \leftarrow x + \mathrm{Attn}(\mathrm{LN}_1(x)), \qquad x \leftarrow x + \mathrm{MLP}(\mathrm{LN}_2(x)) \) one nanoGPT block, as in Block.forward.

Random 32×32 linear layers, computed now. Not a trained model.

4 · The components

C5 · LayerNorm, and where you put it

LayerNorm rescales each token's vector, so sub-layers see inputs of one size at any depth: steadier training.

\( \mathrm{LN}(x) = \gamma \odot \dfrac{x-\mu}{\sqrt{\sigma^2+\epsilon}} + \beta \) for one token's vector \(x\) of \(d\) numbers: \(\mu\), \(\sigma^2\) are their mean and variance; \(\gamma\), \(\beta\) are learned.

Post-LN: the 2017 Transformer. Pre-LN: GPT-2 and nanoGPT.

4 · The components

C6 · Output scores and shared embedding weights

The last vector at each position becomes one score per vocabulary token, reusing the token table E.

\( p(x_{t+1} \mid x_{\le t}) = \mathrm{softmax}\big(\mathrm{LN}_f(h_t)\,E^\top\big) \) \(h_t\): the last block's output at position \(t\). \(E\): the token table from C1, reused.

Counts follow nanoGPT's model.py. GPT-2 small: 124,439,808.

4 · The components, in nanoGPT · summary

What makes Transformers (GPT) great?

A compound system built on 70 years of work: six simple parts, each with one job.

ComponentWhat it gives the modelWhere it comes from
C1EmbeddingsA vector per token, plus one per positionWords become numbers; order is keptWord vectors, NPLM (2003)
C2Causal attentionEach token looks at every earlier token in one stepLong-range context; all positions train at onceAttention (2015); self-attention (2017)
C3MLPWorks on each token on its ownMost of the parameters, to store what it learnsMultilayer perceptron, backprop (1986)
C4ResidualsAdds each layer’s result to its inputDeep stacks still trainResNet (2016)
C5LayerNormPuts every token’s vector on a standard scaleSteadier training as the stack gets deepLayer normalization (2016)
C6Output headLast vector → a score per vocabulary tokenA probability for the next tokenSoftmax, NPLM (2003)
+Next-token lossPredict the token after every positionAny text is training dataShannon (1948)

4 · The components

What nanoGPT does not have — yet

LLaMA and DeepSeek-V3 keep nanoGPT's layout and change these parts.

What changesWhy it mattersUsed in
RoPE C1rotates queries and keys by positionscores depend on relative distanceLLaMA, DeepSeek-V3
RMSNorm C5rescales the vector without subtracting its meansimilar performance, 7–64% faster in its paperLLaMA, DeepSeek-V3
SwiGLU C3a gated MLP, width 8/3·dbeats ReLU and GELU at equal sizeLLaMA
GQA C2query heads share key/value headsless memory for stored keys and valuesLlama 2 34B, 70B
MoE C3many MLPs; a few run per tokenmore parameters, same computeDeepSeek-V3: 671B, 37B active

Know what each name refers to for now. We revisit these in Session 02.

BREAK 2
10 minutesBack at 20:00

Coming up in Part C: we remove these components from a small GPT, one at a time, and you predict what breaks.

Part C

Break a GPT

ConvNeXt added Transformer ideas to a ConvNet, one at a time. Now we take parts out of a small GPT, one at a time. Then you do it.

5 · Ablating a small GPT

What we will break: a small GPT with nanoGPT’s structure

nanoGPT is a minimal GPT: a model file of about 300 lines with the six parts from Part B. Ours is a one-file copy, with the model and the data made small enough that one run takes about a minute on a laptop.

The six parts: C1 embeddings · C2 causal self-attention · C3 MLP with GELU · C4 residuals · C5 LayerNorm · C6 output head

Data · TinyShakespeare (char-rnn)
First Citizen:
Before we proceed any further, hear me speak.

All:
Speak, speak.

1,115,394 characters (1.1 MB), one token each. 65 distinct: A–Z, a–z, space, newline and !$&',-.3:;?. First 90% to train, last 10% to validate.

Model
3.3 MB

818,241 parameters × 4 bytes (float32)

For scale: GPT-3 has 175 billion parameters, 350 GB at 2 bytes each.

4 blocks, 4 heads, width 128, context 64 characters.

Training · our runs

3,000 steps of 32 sequences; AdamW, learning rate 3 × 10⁻⁴; gradients clipped at 1.0. Curve: mean of 3 seeds, laptop CPU.

All of this is held fixed for every run. Curve from our runs of ablation_base.py, 29 Sep 2026: training loss every 10 steps.

5 · Ablating a small GPT

The setup

Held fixed, for every run

Everythingthe data, model and training on the previous slide

Seeds1337, 1338, 1339

Measurevalidation loss, averaged over 20 batches

Changed, one at a time

Remove LayerNorm

Remove residual connections

Remove position embeddings

Four heads → one head

python -m venv gpt && source gpt/bin/activate   # optional: a clean environment
pip install torch                               # Python 3.9+; no GPU needed
curl -O https://costa-nus.github.io/CEG5305_GenAI/part2/lecture-1/code/ablation_base.py
python ablation_base.py --out baseline.json     # 3 seeds; downloads the data
# your change: add one entry to VARIANTS in ablation_base.py, e.g. "deeper": {"n_layer": 12}
python ablation_base.py --variant deeper --out mine.json

ablation_base.py · or in Colab, T4 GPU, about 35 s per run · on a laptop CPU, under 0.5 GB of RAM, about 75 s per run

If anything else differs between two runs, you cannot tell which change caused the difference.

5 · Ablating a small GPT

Predict before you look

Rank the four changes from most damaging to least. Record your prediction before viewing the results.

5 · Ablating a small GPT

All five, side by side

VariantVal lossΔ vs baselineRange, 3 seeds
baseline1.759—1.744–1.768
− LayerNorm1.732−0.0271.724–1.738
− residual connections2.709+0.9502.387–3.337
− positional embeddings1.840+0.0811.798–1.874
single head1.750−0.0091.736–1.765

Mean of 3 seeds, 4 layers, 3,000 steps. Lower is better: green, loss down; red, loss up.

Our runs, 27 Sep 2026, Colab T4 GPU: the model in ablation_base.py. Rerun on a laptop CPU on 29 Sep: the same to within 0.002, except − residual (2.27–2.37). Data: TinyShakespeare (char-rnn).

5 · Ablating a small GPT

How to read an ablation honestly

  1. Train each version more than once. Use a different random seed each time. One run can be lucky or unlucky, so report the average and the range.
  2. Train every version the same way. Same number of steps, same data. Otherwise a difference may come from the training, not from the change.
  3. Say what else the change removed. Taking out a part can also take out parameters, so the model is smaller as well as different.
  4. Report changes that made no difference. “One attention head instead of four: no measurable change” is a real finding.

6 · Homework · for practice

Homework: test one component of the small GPT

An ablation tests what a component contributes by comparing models trained with and without it.

  1. Run the baseline. ablation_base.py as given, 3 seeds.
  2. Change one component, nothing else. Remove it or replace it. Same data, steps and seeds.
  3. Compare and explain. Is the change in validation loss larger than the spread across seeds? What does the component do that explains it?
What you produce

Your code change: a diff against ablation_base.py

A table: validation loss, mean and range over 3 seeds, for the baseline and your version

One plot: the training-loss curves of both

A few sentences: what the component does, why the loss changed or did not, and whether you predicted it

This is how Liu et al. (2022) built ConvNeXt, and how we made tonight's table: one change at a time.

6 · Homework · for practice

Which component to change

Choose one

One of tonight's four, in more depth: does removing LayerNorm still help at 12 layers?

Another: MLP width, depth, GELU → ReLU, tied vs untied embeddings

Or add one nanoGPT lacks: RMSNorm, RoPE

Check before you conclude

Same data, steps and seeds as the baseline

Say what else changed, such as the parameter count

No difference is a result: report it

Loss at step 0 near ln 65 = 4.17; if not, fix that first

This homework is optional practice and carries no course credit.

7 · What comes next

Session 02 — how the model is actually trained

Training on large datasets

Minimising the next-token prediction loss over trillions of tokens.

Pre-training data

Where tokens come from, and why most are thrown away.

Scaling laws

Predicting that loss from model size, data and compute, before spending the compute.

Adapting with fewer trainable parameters

LoRA and parameter-efficient fine-tuning.

Where we got to

Why do LLMs work so well?
Architecture, or scale?

Answered today
one hop, all at once base−partloss ↑

Direct connections between tokens and parallel training let Transformers handle larger workloads. Removing components reveals how each affects prediction loss.

Still open
size · data · compute ? loss ↓ by how much

How much larger models, more data and more compute improve performance. We examine this in Session 02.

Questions.

← → or PageUp/PageDown · F full screen · O overview · B black screen