Session 01 of 2 · Architecture (session 02: Training)
CEG5305 Introduction to Generative AI with Foundation Models
Junyuan Hong · AY2026/27 Semester 1
Recap · Part 1, with Liu Xingyu
| Lecture | Topic | Key message |
|---|---|---|
| 01 | Generative and autoregressive models | A generative model learns p(x); the chain rule turns it into next-token predictions, and a Transformer with a causal mask computes them |
| 02 | Learning by maximum likelihood | Training minimises the KL divergence to the data, which means minimising the cross-entropy of each true next token |
| 03 | Latent variable models, the VAE | Explain data with hidden variables z; the ELBO makes training possible |
| 04 | Diffusion models | Add noise step by step, then learn to remove it: DDPM, and DiT and Stable Diffusion in practice |
| 05 | Normalizing flows, flow matching | Invertible maps give an exact likelihood; flow matching learns the path from noise to data directly |
Today builds on 01 and 02: the chain rule, the causal mask and the cross-entropy loss all come back.
Looking ahead · Part 2, with Junyuan Hong
| Lecture | Topic | Key message |
|---|---|---|
| 1 · 01 | Architecture of LLMs today | What a GPT is made of, and how to find out which parts matter |
| 1 · 02 | Training, scaling laws, LoRA | A fixed compute budget: how much data, how big a model; then adapting a trained model cheaply |
| 2 · 01 | Post-training: instruction tuning, RLHF | A base model continues text; post-training teaches it to answer |
| 2 · 02 | Prompting, in-context learning, chain of thought | Changing what a model does without changing a single weight, and where that stops working |
| 3 | Self-supervised foundation models | Learning from unlabelled data by predicting one part from another, beyond text |
| 4 | Large multi-modal models | Connecting a vision encoder to an LLM so that it can see |
From what the model is, to how it is trained, aligned, prompted, and extended beyond text.
Outline · Lecture 1 = Topic 4, six hours
Training on large datasets · pre-training data · scaling laws · LoRA fine-tuning
Today's question
And the part of that question we can test today:
Is it the architecture, or the scale?
Today we test model design. Session 02 examines model size, data and compute.
Today's question · the record so far
Scores from the papers' tables (Vaswani T2, Devlin T1, Touvron 2023a T9 and 2023b T19, DeepSeek-AI T3, OpenAI T2); compute of the MMLU models from Epoch AI, mostly estimates. *GPT-4: estimated compute (bar: 90% interval), post-trained score. Logos: Simple Icons.
Part A
Compare the models’ outputs, then see how their designs changed.
1 · Show
Counts from Alice in Wonderland.
Used in: Google's translator in 2007.
“The Eiffel Tower is”
Suggest a prompt for all three.
1 · Show · demo 1
1 · Show · demo 2
1 · Show · demo 3
1 · Show
the eiffel tower is such a large letters it mouse to think about this morning i could the garden and beg your pardon cried alice notThe Eiffel Tower isself cats court of the side, who was going to the arch-arching of tell, which handed the trembled on ord her. "That's are the beginnous way and she coThe Eiffel Tower is now an international landmark on the Eiffel Tower estate, where it has been restored to its historic glory. It was originally built in 1923 by the British builder, Sir Winston Churchill. …Pre-run on 29 Sep 2026 with the prompt “The Eiffel Tower is”, the first run of each, not chosen for quality. n-gram: word bigram, 22 words. LSTM: temperature 0.7, 150 characters. GPT-2 small: temperature 0.8, top-k 40, the first two sentences of 50 tokens. All three from the demos on the previous slides.
1 · Show
The Transformer can directly connect distant tokens and process positions in parallel during training.
More training data and more computation.
Both contribute. How much does each explain the differences we observed?
Segment 2 · 25 min
n-gram → NPLM → RNN → LSTM → Transformer
2 · History of generative models · the common principle
Our goal is to generate a sentence \([x_1, \dots, x_T]\), one token at a time.
Chain rule: exact for any sequence.
\[ p(x_1,\dots,x_T) = \prod_{t=1}^{T} p(x_t \mid x_1,\dots,x_{t-1}) \]For text, code, DNA or music, these next-token probabilities determine the probability of a complete sequence. Generate a sequence by sampling one token at a time.
Markov chain: keep only the last \(k\) tokens.
\[ p(x_t \mid x_1,\dots,x_{t-1}) \approx p(x_t \mid x_{t-k},\dots,x_{t-1}) \]With \(k\) as long as the sequence, it is exact again, but there are \(V^k\) contexts to learn.
Markov (1913) counted vowel–consonant pairs in Pushkin's Eugene Onegin; Shannon (1948) used such chains to generate English.
2 · History of generative models
2 · History of generative models
Keep only the last \(n-1\) tokens, then count:
\[ p(x_t \mid x_{1},\dots,x_{t-1}) \approx \frac{\mathrm{count}(x_{t-n+1},\dots,x_{t-1},\textcolor{#ef7c00}{x_t})}{\mathrm{count}(x_{t-n+1},\dots,x_{t-1})} \]The model stores counts from the training text rather than learned neural-network weights.
2 · History of generative models
Counted live from the seven sentences shown. Try the trigram model, then click “dog”, “mat”.
Board · derivation 1 · ~5 min
✋ “the cat sat on the ___” and “the dog sat on the ___”: does the table know these are related?
2 · History of generative models
Word vector: \(C\) is a \(|V| \times m\) table of learned numbers; the vector of word \(w\) is its row, \(C(w)\).
\( x = \big(C(w_{t-n+1}), \dots, C(w_{t-1})\big) \qquad h = \tanh(Hx + d) \qquad P(w_t = i \mid \text{context}) = \mathrm{softmax}(Uh + b)_i \)
Eq. 1 of Bengio et al. (2003), with its optional direct connections \(Wx\) left out. \(C, H, d, U, b\) are trained together by gradient descent. Random, untrained weights here.
2 · History of generative models
A feed-forward network (MLP) reads a fixed window of learned word vectors. Similar words get similar vectors.
2 · History of generative models · how neural models learn
From the NPLM on, weights are trained, not counted: move each weight against its gradient.
loss \( \mathcal{L} = -\log p_\theta(x_{t+1} \mid x_{\le t}) \) update \( \theta \leftarrow \theta - \eta\,\nabla_\theta \mathcal{L} \) backprop \( \dfrac{\partial \mathcal{L}}{\partial h_1} = \dfrac{\partial \mathcal{L}}{\partial h_L}\,\dfrac{\partial h_L}{\partial h_{L-1}} \cdots \dfrac{\partial h_2}{\partial h_1} \)
Loss: Part 1, lecture 02. Backpropagation: Rumelhart et al. (1986). Left: \(L(w) = \tfrac12 (w-1)^2\). Right: one number \(g\) stands for the size of each factor; an illustration, not a trained model. Vanishing gradients in RNNs: Bengio et al. (1994).
2 · History of generative models
\( h_t = \tanh(W\,h_{t-1} + U\,x_t), \qquad y_t = \mathrm{softmax}(V\,h_t) \) \(x_t\): the word vector of token \(t\). \(h_t\): the memory, the same size at every step.
Left: the NPLM sees only its window. Right: a 6-unit tanh RNN with random weights; the bars show ‖∂hₜ/∂x₁‖, how much the first input can still change the memory.
2 · History of generative models
\( h_t = \tanh(W h_{t-1} + U x_t) \): one state, the same weights at every step.
2 · History of generative models
\( c_t = f_t \odot c_{t-1} + i_t \odot g_t, \qquad h_t = o_t \odot \tanh(c_t) \) gates \(f_t, i_t, o_t = \sigma(\cdot)\), between 0 and 1; new content \(g_t = \tanh(\cdot)\); each computed from \(h_{t-1}\) and \(x_t\).
After Graves (2013), Fig. 2. Scalar version of one cell.
2 · History of generative models
Forget gate f held fixed, nothing added. Dashed: an RNN shrinking by 0.6 per step (stylised).
2 · History of generative models · demo
What it generates: the next pen move. At each step the LSTM outputs a probability distribution over how far the pen moves in x and y, and whether it lifts off the page. It samples one move, draws it, and reads it back in.
Graves (2013), §4: pen offsets from a mixture of 2-D Gaussians plus an end-of-stroke probability; §5: a soft window over the text tells it which letter it is writing. Left: the pen path is the Hershey script font, not a model sample; the orange ellipses show the kind of distribution output at each step.
2 · History of generative models
| Still a bottleneck | Consequence |
|---|---|
| ① Step \(t\) waits for step \(t-1\): \(n\) steps in a row, in training too. | Slow training. The \(n\) positions cannot be computed at the same time, so training time grows with sequence length however many GPU cores are free. It is not more arithmetic: per layer, recurrence costs \(O(n \cdot d^2)\) and attention \(O(n^2 \cdot d)\). The difference is \(O(n)\) sequential steps against \(O(1)\). |
| ② Everything about the past must fit in one fixed-size state. | Limited capacity. A long input is squeezed into the same \(d\) numbers, so detail is lost, and a word \(n\) steps back reaches the output only through \(n\) updates. A plain encoder–decoder translated worse as sentences got longer. |
Costs per layer and sequential steps: Vaswani et al. (2017), Table 1. Longer sentences, lower BLEU: Bahdanau et al. (2015), Fig. 2.
2 · History of generative models
Bottleneck ②: the encoder’s last state \(c\) is all the decoder gets (Sutskever et al., 2014). Attention (Bahdanau et al., 2015) lets each decoder step look back at every encoder state.
\( \alpha_{ij} = \mathrm{softmax}_j\big(a(s_{i-1}, h_j)\big), \qquad c_i = \sum_j \alpha_{ij}\, h_j, \qquad s_i = f(s_{i-1}, y_{i-1}, \textcolor{#ef7c00}{c_i}) \) Score each \(h_j\) against \(s_{i-1}\), turn the scores into weights \(\alpha_{ij}\), feed their weighted average \(c_i\) to step \(i\).
Formulas: Bahdanau et al. (2015), §3.1, Eqs. 4–6; \(a\) is a small learned network. Alignment weights in the diagram are hand-set for illustration, not from a trained model.
2 · History of generative models
Simplified RNN: each update retains a fraction g, as in the LSTM example. Attention weights are set by hand.
2 · History of generative models
(1) seq2seq → encoder–decoder (2) recurrence → self-attention
Encoder: reads the whole source at once; every word attends to every other. Word order comes from a position vector added to each word, not from reading left to right.
Decoder: writes the target. Cross-attention reads the encoder, the seq2seq attention without the RNN. Its own self-attention is masked: why, on the next slide.
After Fig. 1 and §3.1 of Vaswani et al. (2017), simplified.
2 · History of generative models
\( \mathrm{softmax}(QK^\top/\sqrt{d} + M)\,V \), with \(M_{ij} = -\infty\) for \(j > i\), else 0: Vaswani et al. (2017), §3.2.3. Weights hand-set; Part B computes it with real numbers.
Board · derivation 2 · ~6 min
Attention compares O(n²) token pairs, but computes those comparisons in parallel. An RNN must process n steps in order.
2 · History of generative models
From a French sentence, predict its English translation.
Predict the next token, at every position.
GPT-1: a “12-layer decoder-only transformer with masked self-attention heads” (Radford et al., 2018). Translation by prompt, “english sentence = french sentence”: GPT-2 5 BLEU on WMT-14 En→Fr (Radford et al., 2019, §3.7); GPT-3 few-shot 32.6 (Brown et al., 2020, Table 3.4).
2 · History of generative models
Decoder-only: nanoGPT model.py. Encoder–decoder: after Fig. 1 of Vaswani et al. (2017).
2 · History of generative models
Each block combines attention and a feed-forward network (MLP). Position vectors encode token order.
2 · History of generative models · demo
2 · History of generative models · demo
2 · History of generative models
| Model | Context | Path length | Parallel training | Main strength or limit |
|---|---|---|---|---|
| n-gram | n−1 tokens | — | n/a | Cannot share similar contexts |
| NPLM | fixed window | 1, in window | yes | Window cannot grow |
| RNN | unbounded* | O(n) | no | Forgets; sequential |
| LSTM | long | O(n) | no | Longer memory; still sequential |
| Transformer | full window | O(1) | yes | Direct access; O(n²) comparisons |
* In principle. Shrinking gradients make learning from distant tokens difficult.
2 · History of generative models
Transformers also process non-text tokens. Each task uses its own input encoding and output layer.
Examples: eagle, USFWS; Armstrong, Apollo 11, NASA; Muybridge (1878); all public domain. DiT sample, facebookresearch/DiT (CC BY-NC 4.0). Ubiquitin, PDB 1UBQ.
Part A — what attention changes
These explain advantages over RNNs. Part B tests how other design choices improve performance.
Coming up in Part B: an experiment that changes one thing at a time, then nanoGPT’s six components.
Part B
Test one design change at a time, then examine nanoGPT’s components.
3 · The ConvNeXt experiment
A Transformer differs from a ConvNet in many ways at once:
If two models differ in many ways, a higher score does not tell us which change helped.
3 · The ConvNeXt experiment
Start with ResNet-50: 76.1% ImageNet accuracy.
Add one design or training change at a time.
Measure accuracy after each change, using similar amounts of computation.
Liu et al. (2022), A ConvNet for the 2020s. Diagram schematic: effect sizes made up; the measured ones are on the next slide.
3 · The ConvNeXt experiment
Liu et al. (2022), Fig. 2 and §2. ImageNet-1K top-1, ResNet-50 → ConvNeXt-T; Swin-T as the target.
3 · The ConvNeXt experiment
Some gains came from changes besides attention.
Training changes alone: +2.7 percentage points in accuracy.
Twelve design changes together: +3.2 more.
That attention is useless.
Whether the same gains apply to language: this experiment tests image classification.
We can use the same experimental method. Change one thing, hold the rest fixed, measure. Part C does it to a GPT.
Segment 4 · 28 min
From the paper to about 300 lines of model.py
4 · The components
Tokens become vectors (C1), pass through n_layer blocks with the same structure (C2–C5), then become next-token probabilities (C6). First, what it is trained to do; then each component.
Board · derivation 3 · ~7 min
All these models learn to predict the next token. Their designs, the six components next, change how they compute \(p_\theta\).
4 · The components
Every position predicts the token after it; the loss averages −log p of the true next token. Training lowers it by gradient descent.
4 · The components
The same word gets a different vector at a different position, because a position vector is added.
\( x_t = E[\mathrm{idx}_t] + P[t] \) \(E\): \(V \times d\) token table. \(P\): \(T \times d\) position table.
4 · The components
A score compares one token’s query vector with another’s key vector. Higher scores give that token more attention.
\( Q = XW_Q,\; K = XW_K,\; V = XW_V, \qquad S = QK^\top/\sqrt{d_k} \) \(X\): one row per token. \(S_{ij}\): query \(i\) against key \(j\).
Five tokens, d = 4, random weights.
4 · The components
Each output averages its own and earlier tokens’ value vectors, using attention weights. This is where information passes between tokens.
\( A = \mathrm{softmax}(S + M), \qquad \mathrm{out} = A\,V \) \(M_{ij} = -\infty\) for \(j > i\), else 0. Each row of \(A\) sums to 1.
4 · The components
Attention mixes tokens; the MLP then updates each token on its own. Its nonlinearity, GELU, is a smooth ReLU whose slope is not zero for negative inputs, so gradients still pass.
\( \mathrm{MLP}(x) = W_2\,\mathrm{GELU}(W_1 x) \) \(W_1\): \(4d \times d\), \(W_2\): \(d \times 4d\). Where are most of the parameters?
4 · The components
Each sub-layer adds its result to x, so gradients flow straight down x instead of shrinking layer by layer.
\( x \leftarrow x + \mathrm{Attn}(\mathrm{LN}_1(x)), \qquad x \leftarrow x + \mathrm{MLP}(\mathrm{LN}_2(x)) \) one nanoGPT block, as in Block.forward.
Random 32×32 linear layers, computed now. Not a trained model.
4 · The components
LayerNorm rescales each token's vector, so sub-layers see inputs of one size at any depth: steadier training.
\( \mathrm{LN}(x) = \gamma \odot \dfrac{x-\mu}{\sqrt{\sigma^2+\epsilon}} + \beta \) for one token's vector \(x\) of \(d\) numbers: \(\mu\), \(\sigma^2\) are their mean and variance; \(\gamma\), \(\beta\) are learned.
Post-LN: the 2017 Transformer. Pre-LN: GPT-2 and nanoGPT.
4 · The components
The last vector at each position becomes one score per vocabulary token, reusing the token table E.
\( p(x_{t+1} \mid x_{\le t}) = \mathrm{softmax}\big(\mathrm{LN}_f(h_t)\,E^\top\big) \) \(h_t\): the last block's output at position \(t\). \(E\): the token table from C1, reused.
Counts follow nanoGPT's model.py. GPT-2 small: 124,439,808.
4 · The components, in nanoGPT · summary
A compound system built on 70 years of work: six simple parts, each with one job.
| Component | What it gives the model | Where it comes from | |
|---|---|---|---|
| C1 | EmbeddingsA vector per token, plus one per position | Words become numbers; order is kept | Word vectors, NPLM (2003) |
| C2 | Causal attentionEach token looks at every earlier token in one step | Long-range context; all positions train at once | Attention (2015); self-attention (2017) |
| C3 | MLPWorks on each token on its own | Most of the parameters, to store what it learns | Multilayer perceptron, backprop (1986) |
| C4 | ResidualsAdds each layer’s result to its input | Deep stacks still train | ResNet (2016) |
| C5 | LayerNormPuts every token’s vector on a standard scale | Steadier training as the stack gets deep | Layer normalization (2016) |
| C6 | Output headLast vector → a score per vocabulary token | A probability for the next token | Softmax, NPLM (2003) |
| + | Next-token lossPredict the token after every position | Any text is training data | Shannon (1948) |
4 · The components
LLaMA and DeepSeek-V3 keep nanoGPT's layout and change these parts.
| What changes | Why it matters | Used in | |
|---|---|---|---|
| RoPE C1 | rotates queries and keys by position | scores depend on relative distance | LLaMA, DeepSeek-V3 |
| RMSNorm C5 | rescales the vector without subtracting its mean | similar performance, 7–64% faster in its paper | LLaMA, DeepSeek-V3 |
| SwiGLU C3 | a gated MLP, width 8/3·d | beats ReLU and GELU at equal size | LLaMA |
| GQA C2 | query heads share key/value heads | less memory for stored keys and values | Llama 2 34B, 70B |
| MoE C3 | many MLPs; a few run per token | more parameters, same compute | DeepSeek-V3: 671B, 37B active |
Know what each name refers to for now. We revisit these in Session 02.
Coming up in Part C: we remove these components from a small GPT, one at a time, and you predict what breaks.
Part C
ConvNeXt added Transformer ideas to a ConvNet, one at a time. Now we take parts out of a small GPT, one at a time. Then you do it.
5 · Ablating a small GPT
nanoGPT is a minimal GPT: a model file of about 300 lines with the six parts from Part B. Ours is a one-file copy, with the model and the data made small enough that one run takes about a minute on a laptop.
The six parts: C1 embeddings · C2 causal self-attention · C3 MLP with GELU · C4 residuals · C5 LayerNorm · C6 output head
First Citizen: Before we proceed any further, hear me speak. All: Speak, speak.
1,115,394 characters (1.1 MB), one token each. 65 distinct: A–Z, a–z, space, newline and !$&',-.3:;?. First 90% to train, last 10% to validate.
818,241 parameters × 4 bytes (float32)
For scale: GPT-3 has 175 billion parameters, 350 GB at 2 bytes each.
4 blocks, 4 heads, width 128, context 64 characters.
3,000 steps of 32 sequences; AdamW, learning rate 3 × 10⁻⁴; gradients clipped at 1.0. Curve: mean of 3 seeds, laptop CPU.
All of this is held fixed for every run. Curve from our runs of ablation_base.py, 29 Sep 2026: training loss every 10 steps.
5 · Ablating a small GPT
Everythingthe data, model and training on the previous slide
Seeds1337, 1338, 1339
Measurevalidation loss, averaged over 20 batches
Remove LayerNorm
Remove residual connections
Remove position embeddings
Four heads → one head
python -m venv gpt && source gpt/bin/activate # optional: a clean environment
pip install torch # Python 3.9+; no GPU needed
curl -O https://costa-nus.github.io/CEG5305_GenAI/part2/lecture-1/code/ablation_base.py
python ablation_base.py --out baseline.json # 3 seeds; downloads the data
# your change: add one entry to VARIANTS in ablation_base.py, e.g. "deeper": {"n_layer": 12}
python ablation_base.py --variant deeper --out mine.json
ablation_base.py · or in Colab, T4 GPU, about 35 s per run · on a laptop CPU, under 0.5 GB of RAM, about 75 s per run
If anything else differs between two runs, you cannot tell which change caused the difference.
5 · Ablating a small GPT
Rank the four changes from most damaging to least. Record your prediction before viewing the results.
5 · Ablating a small GPT
| Variant | Val loss | Δ vs baseline | Range, 3 seeds |
|---|---|---|---|
| baseline | 1.759 | — | 1.744–1.768 |
| − LayerNorm | 1.732 | −0.027 | 1.724–1.738 |
| − residual connections | 2.709 | +0.950 | 2.387–3.337 |
| − positional embeddings | 1.840 | +0.081 | 1.798–1.874 |
| single head | 1.750 | −0.009 | 1.736–1.765 |
Mean of 3 seeds, 4 layers, 3,000 steps. Lower is better: green, loss down; red, loss up.
Our runs, 27 Sep 2026, Colab T4 GPU: the model in ablation_base.py. Rerun on a laptop CPU on 29 Sep: the same to within 0.002, except − residual (2.27–2.37). Data: TinyShakespeare (char-rnn).
5 · Ablating a small GPT
6 · Homework · for practice
An ablation tests what a component contributes by comparing models trained with and without it.
Your code change: a diff against ablation_base.py
A table: validation loss, mean and range over 3 seeds, for the baseline and your version
One plot: the training-loss curves of both
A few sentences: what the component does, why the loss changed or did not, and whether you predicted it
This is how Liu et al. (2022) built ConvNeXt, and how we made tonight's table: one change at a time.
6 · Homework · for practice
One of tonight's four, in more depth: does removing LayerNorm still help at 12 layers?
Another: MLP width, depth, GELU → ReLU, tied vs untied embeddings
Same data, steps and seeds as the baseline
Say what else changed, such as the parameter count
No difference is a result: report it
Loss at step 0 near ln 65 = 4.17; if not, fix that first
This homework is optional practice and carries no course credit.
7 · What comes next
Minimising the next-token prediction loss over trillions of tokens.
Where tokens come from, and why most are thrown away.
Predicting that loss from model size, data and compute, before spending the compute.
LoRA and parameter-efficient fine-tuning.
Where we got to
Direct connections between tokens and parallel training let Transformers handle larger workloads. Removing components reveals how each affects prediction loss.
How much larger models, more data and more compute improve performance. We examine this in Session 02.
Questions.